GPT-5.6 Sol in Ultrafast Mode—and Why Voice Assistants and Developers Are Waiting for It

GPT-5.6 Sol in Ultrafast Mode—and Why Voice Assistants and Developers Are Waiting for It

Ondřej Barták
Ondřej Barták
Entrepreneur and Programmer
19. 8. 2026
3 minutes reading
Listen to the article
Audio version of the article
GPT-5.6 Sol in Ultrafast Mode—and Why Voice Assistants and Developers Are Waiting for It

OpenAI has introduced a new service tier called Ultrafast, in which its GPT-5.6 Sol model runs up to fourteen times faster than with standard processing. The company promises speeds of up to 750 output tokens per second, with the entire technology built on Cerebras hardware. For now, this is limited early access for a select group of API customers, not a feature that anyone can enable directly in the ChatGPT interface. 

Until now, there has been an unpleasant trade-off. When a company needed an immediate response, it had to use a smaller or narrowly specialized model. While it was fast, it often fell short when dealing with more complex problems, working with multiple tools, or handling longer chains of decisions. A more powerful model, on the other hand, took so long to think that it was unsuitable for live interaction. Ultrafast is intended to eliminate this trade-off. OpenAI describes it with the slogan “more useful work per second” and is targeting situations where results are needed within seconds rather than minutes.

Who Ultrafast is for

OpenAI highlighted several use cases where it is already seeing initial results. During an outage of a critical system, the model can read operational logs, review recent code changes and technician reports, and look for the likely cause while the incident is still ongoing. In the financial sector, use cases include analyzing market signals, assessing transactions, and detecting suspicious activity as conditions continually change.

Voice and customer support are major areas of focus. When an assistant needs to retrieve data from three systems, verify terms, and only then provide an answer, every delay adds up during the call, and the customer notices immediately. The same applies to e-commerce, where the model answers product questions, checks inventory, and resolves payment issues before the shopper leaves and abandons a full shopping cart.

Research teams at OpenAI routinely run a set of experiments overnight and review the results in the morning. With greater speed, this development cycle becomes short enough to test several different variants in succession during a single working day. 

Collaboration with Cerebras

Ultrafast technology is based on OpenAI’s ongoing collaboration with Cerebras, which manufactures so-called wafer-scale processors. Put simply, these are exceptionally large chips designed specifically for artificial intelligence computing. Instead of dividing a computational task among a large number of smaller units, as much of the work as possible is kept together in one place.

This approach is particularly well suited to large language models. Each successive token builds on the previous one, so the resulting speed depends entirely on how short the pause is between individual steps. Once this delay is reduced, text stops arriving in separate batches and begins to flow smoothly.

One crucial detail can easily be overlooked. The published figure describes the output generation rate, not the time it takes to receive the complete response. That still includes network speed, prompt length, calls to external tools, and waiting for surrounding systems. A business application therefore will not speed up automatically simply by switching to a different mode.

Pricing and availability

There is currently no official price list, billing method, or usage limits. It is not publicly known under what conditions the company achieves the stated 750 tokens per second, or how the speed will be affected by longer requests and concurrent use by multiple users. OpenAI has not yet announced a launch date for Europe. It also notes that the availability of API speed tiers in individual countries depends on local legal requirements. Access is therefore currently limited to a small group of companies testing the Sol model in this mode for programming, trading, financial analysis, and customer support. Interested parties can apply using the form and wait for OpenAI to release additional capacity. 

Category:AI
Did you enjoy this article?
Discover more interesting posts on our blog
Back to blog

Related posts

Public Attitudes Toward AI in 2026: Asia Is Enthusiastic, Europe and America AnxiousPublic Attitudes Toward AI in 2026: Asia Is Enthusiastic, Europe and America Anxious
New data shows that AI is dividing the world: enthusiasm prevails in Asia, while Europe and America are gripped by anxiety. Even so, people expect it to have an ever-greater impact on life and work.
6 min read
21. 8. 2026
AI in IT Support Saves Hours, Yet Workloads Have IncreasedAI in IT Support Saves Hours, Yet Workloads Have Increased
AI in IT support demonstrably saves time, but also creates more work maintaining, monitoring, and fine-tuning systems. What matters is whether companies measure activity or actual impact.
3 min read
21. 8. 2026
Stripe Tells Investors It Is Buying OpenRouter and the Singularity Has BegunStripe Tells Investors It Is Buying OpenRouter and the Singularity Has Begun
Stripe is acquiring OpenRouter for billions to accelerate AI deployment. It also surprised investors by claiming that the technological singularity has already begun.
5 min read
21. 8. 2026
Přihlaste se k odběru našeho newsletteru
Zůstaňte informováni o nejnovějších příspěvcích, exkluzivních nabídkách, a aktualizacích.
CodedTrip

Operated by CodedTrip LLC, USA.

YouTube
TikTok