On July 20, news emerged that Google is in the process of developing a server chip, internally codenamed 'Frozen v2'. This chip is set to cement the foundational computing architecture of the Gemini large-scale model into hardware, thereby achieving a reduction in both energy consumption and latency during AI inference processes. It is anticipated that the number of tokens processed per unit of power consumption will soar to 6 to 10 times that of current chips, with a potential deployment timeline as early as 2028.
