Details regarding the Ultrafast API have surfaced within OpenAI's platform and API documentation, hinting at the company's potential intention to broaden the accessibility of this API. Presently, OpenAI is in the process of developing three distinct processing speed alternatives for the Responses API Playground: Standard, Fast, and Ultrafast. However, this functionality remains concealed for the time being. In the past, OpenAI provided a sneak peek of the Ultrafast mode, which is powered by GPT-5.6 Sol. This mode is capable of generating outputs at a remarkable rate of up to 750 tokens per second, representing a substantial 14-fold surge in inference speed when compared to the Standard mode. Nevertheless, at present, this mode is exclusively accessible to a select group of customers and might witness further expansion during the DevDay scheduled for September 29. Regarding whether the already-released GPT-6 series models are compatible with the Ultrafast mode, official confirmation is still pending.
