Google is in the process of developing a server chip, internally dubbed 'Frozen v2', with the primary objective of significantly enhancing the operational efficiency of the Gemini model. This innovative chip is set to integrate the model's foundational computing architecture directly into hardware, thereby achieving a reduction in both energy consumption and latency during AI inference processes. Notably, it is anticipated that the number of tokens processed per unit of power consumption will soar to a level 6 to 10 times higher than that of Google's current in-house AI chips. The project is slated for deployment as early as 2028, albeit on a limited initial production scale. It is envisioned as a technological experimental platform, designed to complement and enhance the capabilities of Google's existing Tensor Processing Units (TPUs).
