OpenAI's first self-developed AI inference chip, Jalapeño, has demonstrated comprehensive performance superiority over Nvidia's GB200 and GB300 in real-world tests. In evaluations using three models—GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T—Jalapeño achieved 1.5 to 1.9 times the AI workload per watt compared to competing systems, with end-to-end latency reduced to just 28% to 59%. The chip has a rated power of 700 watts, with actual sustained power consumption at or below 550 watts. It is planned for small-scale deployment by the end of this year, with gradual scale-up expected by 2027.
