Over the past several months, OpenAI and Anthropic have been in a relentless pursuit of enhancing their flagship models. While Google's Flash series has undergone frequent iterations, its Pro lineup has maintained a notable silence. Recently, a model dubbed gemini-3.8-flash surfaced in a large-scale model blind test arena, showcasing performance significantly superior to that of the official Gemini 3.8 Flash version. This has sparked speculation that it could be Google's next-generation flagship, the Gemini 4 Pro.
Practical tests conducted by developers have revealed that this model exhibits exceptional proficiency in SVG drawing, 3D scene construction, web design, and game development. Some developers even claim to have unearthed internal terminal information that is suspected to be linked to this model. Concurrently, a widely circulated benchmark comparison table indicates that the model outperforms GPT-6 Astra and Claude Fable across multiple tests. However, none of this information has received official confirmation, and at present, only the model's outstanding performance can be verified.
The rumors also extend to RSI (Recursive Self-Improvement). Google is intensifying AI's involvement in model development, having already assigned certain researcher tasks to Agents and officially mentioning the Agent-based recursive enhancement of models. Related research papers also delve into the topic of RSI. Anthropic and China's Zhipu are similarly advancing AI-driven research and development efforts.
Previously, Amodei urged frontier AI companies to exercise caution and slow down, particularly highlighting the risks associated with RSI. Several labs have acknowledged these risks, but due to fierce competition, no unified regulations have been established. Traces of testing for next-generation models have already emerged at Google, Anthropic, and OpenAI. The call to slow down has not impeded model development; rather, it has merely shed light on the inherent risks.
