OpenAI has introduced a new service tier called Ultrafast, in which its GPT-5.6 Sol model runs up to fourteen times faster than with standard processing. The company promises speeds of up to 750 output tokens per second, with the entire technology built on Cerebras hardware. For now, this is limited early access for a select group of API customers, not a feature that anyone can enable directly in the ChatGPT interface.
Until now, there has been an unpleasant trade-off. When a company needed an immediate response, it had to use a smaller or narrowly specialized model. While it was fast, it often fell short when dealing with more complex problems, working with multiple tools, or handling longer chains of decisions. A more powerful model, on the other hand, took so long to think that it was unsuitable for live interaction. Ultrafast is intended to eliminate this trade-off. OpenAI describes it with the slogan “more useful work per second” and is targeting situations where results are needed within seconds rather than minutes.
Who Ultrafast is for
OpenAI highlighted several use cases where it is already seeing initial results. During an outage of a critical system, the model can read operational logs, review recent code changes and technician reports, and look for the likely cause while the incident is still ongoing. In the financial sector, use cases include analyzing market signals, assessing transactions, and detecting suspicious activity as conditions continually change.
Voice and customer support are major areas of focus. When an assistant needs to retrieve data from three systems, verify terms, and only then provide an answer, every delay adds up during the call, and the customer notices immediately. The same applies to e-commerce, where the model answers product questions, checks inventory, and resolves payment issues before the shopper leaves and abandons a full shopping cart.
Research teams at OpenAI routinely run a set of experiments overnight and review the results in the morning. With greater speed, this development cycle becomes short enough to test several different variants in succession during a single working day.
Collaboration with Cerebras
Ultrafast technology is based on OpenAI’s ongoing collaboration with Cerebras, which manufactures so-called wafer-scale processors. Put simply, these are exceptionally large chips designed specifically for artificial intelligence computing. Instead of dividing a computational task among a large number of smaller units, as much of the work as possible is kept together in one place.
This approach is particularly well suited to large language models. Each successive token builds on the previous one, so the resulting speed depends entirely on how short the pause is between individual steps. Once this delay is reduced, text stops arriving in separate batches and begins to flow smoothly.
One crucial detail can easily be overlooked. The published figure describes the output generation rate, not the time it takes to receive the complete response. That still includes network speed, prompt length, calls to external tools, and waiting for surrounding systems. A business application therefore will not speed up automatically simply by switching to a different mode.
Pricing and availability
There is currently no official price list, billing method, or usage limits. It is not publicly known under what conditions the company achieves the stated 750 tokens per second, or how the speed will be affected by longer requests and concurrent use by multiple users. OpenAI has not yet announced a launch date for Europe. It also notes that the availability of API speed tiers in individual countries depends on local legal requirements. Access is therefore currently limited to a small group of companies testing the Sol model in this mode for programming, trading, financial analysis, and customer support. Interested parties can apply using the form and wait for OpenAI to release additional capacity.



