NVIDIA's New-Generation Computing Platform Completes the Puzzle, Reducing Token Costs by 35 Times When Combined with DeepSeek Model
8 hour ago / Read about 0 minute
Author:小编   

On August 24 (local time), NVIDIA announced that its inference accelerator, the Groq 3 LPX rack, has entered full-scale mass production. This rack will be deployed alongside the Vera central processing unit and Rubin graphics processing unit in the data centers of the new cloud service provider, Nebius, with plans to officially go live later this year. Additionally, NVIDIA unveiled new test results for the Vera Rubin NVL72 under real-world agent workloads. The tests were based on SemiAnalysis's AgentX workload and utilized the DeepSeek V4 Pro model. The results showed that its throughput per megawatt can be up to 30 times higher than that of the previous-generation GB300 NVL72, while reducing the cost per Token by up to 35 times. This achievement will enable operators to support higher-capacity interactive agents within fixed power and infrastructure budgets or provide services of the same capacity at lower costs.

  • C114 Communication Network
  • Communication Home
7 X 24 Track global technological trends
Hot Topic