OpenAI’s New Ultrafast Mode Makes GPT-5.6 Sol 14x Faster
By Veer Solanki · · 679 words
Topics: AI, AI Agents, AI Computing, AI Inference, AI Infrastructure, AI Performance, Anthropic, Artificial Intelligence
OpenAI announced the upcoming release of the preview capability of GPT-5.6 Sol called Ultrafast, which will be able to perform inference 14x faster than the standard mode, achieving 750 tokens per second. The company is signalling to the industry that it is focusing on making its models perform enterprise-scale operations and not just on making them more capable.

Ultrafast works on chips supplied by Cerebras, which is positioning itself as a provider of solutions for organisations searching to make large-scale applications using LLMs feasible. The applications for ultra-fast inference capabilities include general operations concerning incident response, customer service, financial markets, and e-commerce. This is not some far-fetched application; these are the use cases for which companies are willing to pay substantial sums to employ an LLM.
The obvious question is what is sacrificed for such a substantial gain in inference performance. OpenAI did not comment on this publicly, but it is assumed that some amount of quality is lost due to the lower amount of computation per token. It is unclear what exactly is sacrificed, but the change in the public narrative implies that the company acknowledges a new reality where performance is more important than ever.
However, this assumption is based on evidence that enterprises are choosing between various LLMs based on their ability to deliver more tokens per second, and therefore, more products per second. Enterprises do not have the luxury to choose between a deep but slow model and a shallow but fast model because customers need solutions to their problems in real-time. Anthropic’s Claude fast mode is an example of this type of offering, but it is not as attractive as the 14x gain on offer from OpenAI.
A similar trend is evident at the enterprise level. The fundamental issue with inference at the enterprise level concerns the ability to process more data within a shorter time. If one can afford to process a thousand customer service queries concurrently at a relatively low cost, there is no need to wait for the flagship model to think deeper about the resolution to one query. The same goes for any other application area, from financial markets to incident response.
It is unclear whether GPT-5.6 Sol will be able to sacrifice some of its reasoning capabilities for the sake of ultra-fast inference. The model is already reasoning-centric, and there is no indication that such a design approach is conducive to achieving first-mover status across multiple verticals. The preview capability is available only to “a small group of customers” with promises of wider availability “as capacity grows.” Perhaps there was simply not enough capacity at the moment, but more likely than not, the company is not comfortable admitting that the primary value proposition of GPT-5.6 Sol is its enhanced reasoning capability, as it appears to have difficulty convincing enterprise customers about this.
Anthropic has been emphasising reasoning as well, which implies that if reasoning is no longer the primary value proposition, it is a loss for everyone involved. However, in the new world where all models are more or less equal in terms of their core competencies, the slightest advantage in one particular area can be sufficient for creating a moat. If customers are willing to accept somewhat lower levels of reasoning for substantially better performance, then they will have a harder time choosing between Claude and GPT-5.6 Sol. On the other hand, OpenAI’s ability to attract enterprise customers will be enhanced, as its product offering better meets their needs. The announcement reflects the growing importance of performance as a differentiator in the competitive LLM market landscape.
The infrastructure layer is also critical to consider, as both Cerebras and inference efficiency are integral to the equation. The model capability maturity curve appears to be flattening, and the next battleground is the infrastructure layer. The ability to make GPT-5.6 Sol ultra-fast is essential for the financial viability of applications using the model, as they will benefit from significantly lower computational costs. However, it is still unclear whether this will happen, as there has been no public commentary on this yet.