On August 13, 2026, OpenAI introduced a new Ultrafast Mode for its most powerful language model GPT-5.6 Sol, which accelerates processing by up to 14 times compared to standard operation. The announcement marks a technological step that makes high-performance AI models available at real-time speed for the first time without quality compromises.
14x Acceleration with Full Model Quality
Ultrafast Mode processes GPT-5.6 Sol up to 14 times faster than standard execution, achieving an output throughput of up to 750 output tokens per second. OpenAI emphasizes that this represents "more useful work per second" – not just accelerated processing of identical tasks. According to Cerebras, the hardware partner for infrastructure, the speed increase is achieved without quality losses in full GPT-5.6 Sol intelligence.
The technical foundation is a partnership between OpenAI and chipmaker Cerebras. Cerebras provides the specialized hardware infrastructure that enables this performance boost. The company specializes in AI accelerators and uses wafer-scale integration – a chip architecture in which a single silicon wafer functions as a unified processor.
Limited Preview Phase for API Customers
At launch, Ultrafast is in a limited preview phase. OpenAI initially grants access to only a small group of customers via the OpenAI API. Availability in ChatGPT or other end-user products was not announced. OpenAI stated that access will be expanded gradually as capacity grows.
The company identifies several use cases for Ultrafast: incident response, customer support, financial market analysis, and e-commerce are among the primary deployment areas. All these fields share a need for time-critical processing, where delays of seconds can have business impact.
Positioning Against Competition
OpenAI notes that competitors like Anthropic already offer accelerated versions of their models. Anthropic offers a Fast Mode for Claude. OpenAI emphasizes that its own Ultrafast solution delivers higher speeds than competing alternatives, without providing specific comparison values.
The company frames Ultrafast as a strategic shift: "Until now, true real-time speed typically meant choosing a smaller or specialized model. Ultrafast points in a new direction: more useful work per second." This statement addresses the previous dilemma between model quality and response speed – a tradeoff developers had to make when choosing between powerful and fast models.
Open Questions on Pricing
OpenAI provided no information on pricing for Ultrafast Mode. In community discussions on Reddit, users speculated that the mode could cost roughly 1.5 times as much as standard processing. This estimate was classified as "pretty close," without official confirmation. Some discussants expressed expectations for future specialized hardware for high-speed inference.
The announcement came one day before publication deadline. Whether and when Ultrafast will be available to broader customer groups or as part of ChatGPT remains open.
