On July 21, reports surfaced indicating that NVIDIA's latest-generation Vera Rubin chip architecture has commenced delivery to customers, marking its entry into full-scale mass production. Major AI enterprises are actively deploying systems based on this architecture. NVIDIA Vice President Ian Buck revealed that computing systems leveraging Vera Rubin technology have already been delivered to key customers, such as OpenAI and Anthropic, and are currently undergoing deployment.
The Vera Rubin architecture is tailored specifically for large-scale AI inference and agent computing, offering a remarkable performance boost—up to ten times that of the previous generation. It achieves a tenfold increase in inference throughput per watt, while simultaneously slashing the token cost of AI inference to just one-tenth of the previous generation's level. This architecture seamlessly integrates six self-developed chips, including the Vera CPU and Rubin GPU. It adopts a 100% liquid cooling solution, supports 45℃ hot water cooling, and features a cable-free modular design, significantly reducing the time required for rack installation.
Among the first batch of customers deploying Vera Rubin are hyperscale cloud providers like AWS, Google Cloud, and Microsoft Azure, as well as NVIDIA cloud partners such as CoreWeave and Lambda. Pharmaceutical powerhouse Bristol-Myers Squibb (BMS) has also announced its status as the world's first biopharmaceutical company to adopt the Vera Rubin architecture. BMS has procured the DGX SuperPOD supercomputer, equipped with the Vera Rubin NVL72 system, to bolster AI applications in drug discovery and development, with an anticipated 50% reduction in the drug development cycle.
