Kaiming He’s team has introduced the groundbreaking VISTA system, effectively overcoming the challenge of large models underperforming in the ARC-AGI-3 test. Traditionally, test organizers provided large models solely with a numerical table aligned with pixels, constraining their visual perception and leading to subpar results even for leading large models. The VISTA system revolutionizes this by transforming the numerical table into intuitive images, significantly cutting down on computational power usage while dramatically boosting model scores.
Moreover, the system incorporates a lossless visual memory 'photo album' mechanism. This innovative feature enables the model to recall details from any frame at any point, without it counting towards the total number of operational steps. Paired with concise rule note assistance, VISTA facilitates a remarkable leap in performance. This is achieved without the need for extra model training or the laborious task of crafting thousands of lines of code to develop simulators for each game, a requirement in previous mainstream methods.
With VISTA’s support, Claude Opus 5 achieved a flawless score across all 25 public games in the ARC-AGI-3, completing them in just 0.43 times the number of steps humans took on their initial attempt. GPT-5.6 Sol also impressed with a near-perfect 99 points. The 320B open-source model, which initially scored a mere 1.89 points, witnessed a remarkable surge to 66.93 points after integrating VISTA.
The VISTA system showcases remarkable adaptability, excelling in diverse settings such as 3D environments, web games, and static image comprehension tasks. This string of accomplishments not only confirms Kaiming He’s team's insight that "ARC is a visual problem" but also underscores that enhancing intelligence hinges not just on bolstering the model itself but also on refining external mechanisms for model perception and memory. It paves a fresh, visual pathway for AGI exploration.
