OpenAI announces Ultrafast mode for ChatGPT, promising up to 14 times faster responses

The Ultrafast mode, which OpenAI has opened for preview, aims to generate up to 750 output tokens per second on GPT-5.6 Sol.

12punto

OpenAI has announced Ultrafast, a new service tier aimed at reducing response times in the ChatGPT experience. The mode, developed by the company for GPT-5.6 Sol, aims to offer faster output generation, particularly for complex tasks.

According to the announcement, the Ultrafast mode can provide a speed increase of up to 14 times compared to standard processing workflows. With this performance, it is stated that the model can generate up to 750 output tokens per second. Tokens refer to the basic text fragments that artificial intelligence models use when generating text.

IN PREVIEW STAGE

OpenAI emphasizes that high speed does not only mean shorter waiting times; it can also provide more efficient usage in workflows that require real-time interaction. The company notes that while smaller or more narrowly focused models have typically been preferred for similar speed targets to date, Ultrafast focuses on rapid response generation even in more powerful models.

OpenAI has opened a new high-speed service tier for ChatGPT for preview.

The infrastructure of the Ultrafast mode involves a collaboration with chip manufacturer Cerebras. The new mode is currently in the preview stage and is only accessible to a limited group of customers. OpenAI states that access will be gradually increased as server and infrastructure capacity expands.

The use cases targeted by the company include areas requiring rapid response, such as crisis and incident response, customer service, support systems, financial market analysis, and e-commerce. No clear timeline has been shared regarding when Ultrafast will be opened for general use.