On September 29, during the DevDay presentation, OpenAI revealed the introduction of its Ultra Fast service, encompassing APIs, ChatGPT, and Codex. Through the enhancement of hardware architecture and data scheduling strategies, this service attains an inference speed of up to 750 tokens per second, a remarkable 14 times swifter than the standard mode. This advancement is tailored to cater to the demands of low-latency situations, including real-time interactions and financial transactions.
