OpenAI Launches GPT-Live: Voice AI That Can Do More Than You Think

OpenAI Launches GPT-Live: Voice AI That Can Do More Than You Think

Ondřej Barták
Ondřej Barták
Entrepreneur and Programmer
10. 7. 2026
5 minutes reading
OpenAI Launches GPT-Live: Voice AI That Can Do More Than You Think

OpenAI has introduced GPT-Live, a new family of voice models designed to make conversations with artificial intelligence feel like a real call between two people. The company has already deployed it in ChatGPT Voice and begun making it available to users worldwide. The new technology is based on an architecture that can listen and speak at the same time. This eliminates the choppy pauses users have grown accustomed to with existing voice assistants.

Listening and responding simultaneously

GPT-Live can do both at once. During a call, it signals that it is listening with brief interjections such as “mhmm” or “yeah,” can respond quickly and promptly, and can also remain silent when someone needs a moment to gather their thoughts. Its creators describe the result as a voice environment that is surprisingly easy to talk to.

At the same time, OpenAI says it is their smartest voice model yet. For queries that require web searches, deeper reasoning, or more complex tasks, it passes the work in the background to a more powerful model and returns the finished result to the conversation as soon as it is ready. In the meantime, GPT-Live can continue talking with you without interrupting the conversation. At launch, it uses the GPT-5.5 model in the background, but the company plans to gradually replace this supporting model with newer models in the future.

Why older voice systems were so rigid

To understand what has changed, it is worth recalling how previous generations worked. The first ChatGPT Voice combined three separate models. One transcribed speech into text, another generated a response, and the third converted it back into spoken form. This made it possible to talk to advanced models for the first time, but the complexity came at a cost. Information was lost between the individual components, and responses arrived slowly and unnaturally.

Advanced Voice Mode, which came later, partially solved this issue. It processed and generated audio within a single model, allowing it to respond more quickly and fluidly. However, it still operated one turn at a time. The model had to wait for the user to finish speaking and detected the end of a sentence based on silence. As a result, a brief pause or background noise was enough to make the assistant start speaking at an inappropriate moment.

Two changes that fixed it

GPT-Live addresses these problems through two changes to the model’s architecture. The first is continuous interaction. Instead of processing a sequence of separate messages, it processes input and generates output concurrently. The model therefore decides many times per second what to do next: whether to speak, continue listening, pause, interrupt, or use a tool. In addition to enabling a more natural exchange, this also gives the model a better sense of timing and enables real-time translation.

The second change separated the conversation itself from more demanding work. When a query requires searching, reasoning, or a more autonomous approach, GPT-Live sends the task to another model, such as GPT-5.5. This allows the conversation to continue even while multiple tasks are being processed in the background. The same model architecture also makes it possible to use the latest models at any time and combine their capabilities with fluid conversation.

Millions of people use ChatGPT Voice

More than 150 million people talk to ChatGPT through voice and dictation every week. They use it for everyday hands-free assistance, practicing foreign languages, telling bedtime stories, or simply chatting on the way to work. From now on, tapping the Voice button will give them an experience powered by GPT-Live.

Conversations are intended to feel more like real interactions. You can interrupt the assistant with a question, pause to think, or ask it to slow down. It uses brief interjections to indicate that it is listening. OpenAI has also redesigned nine different voices for GPT-Live.

Users can now control the sophistication of responses to some extent. Three reasoning levels are available: Instant for quick responses, and medium and high for situations where the assistant should spend more time considering its answer. Listening has also improved. When you pause to think, ChatGPT waits instead of interrupting you. If you ask it to be quiet, it complies. And in noisy environments, such as among passing cars or other people’s conversations, it can focus on your voice more effectively.

Visual responses have also been added. Some things are better seen than heard, so during a conversation ChatGPT can display clear cards for weather, stocks, sports, and other topics. Voice also continues to support search, memory, images, and file uploads.

Safer for users

OpenAI emphasizes that it designed GPT-Live to be safe by default. In addition to safety features from its latest models, it added specialized training in sensitive areas and new safeguards created specifically for voice. It expanded testing to include audio-based evaluations and artificially generated recordings that rigorously test the most sensitive topics. These include self-harm, psychosis and mania, emotional dependence on artificial intelligence, violence, and sexual content. The model also underwent so-called red teaming, meaning targeted efforts to identify weaknesses, with a focus on risks unique to voice.

Because voice conversations take place in real time, the company created safeguards that can intervene while the model is speaking. If the system detects a potentially dangerous response, it can direct the model toward a safer alternative, display additional messages or support resources, or, in extreme cases, end the conversation. For topics related to self-harm, OpenAI adapted its support procedures for voice, including offering crisis hotlines verified by experts.

Teenage users have received special safeguards. The company trained age-appropriate behavior directly into the model to reduce the risk of inappropriate responses. Through parental controls, parents can decide whether their teenage child may use ChatGPT Voice, and in higher-risk situations involving signs of possible self-harm or suicidal intent, they may receive an alert. OpenAI also emphasizes that the model is intended for conversation, not voice imitation. It uses a set of predefined voices and safeguards that prevent it from imitating a specific person’s voice.

Availability

The rollout is already underway for ChatGPT users worldwide on iOS, Android, and the ChatGPT.com website. GPT-Live-1 is becoming the default model for Go, Plus, and Pro plans, while users of the free version will use GPT-Live-1 mini. OpenAI plans to make both versions available to developers through the API soon, and interested users can already sign up for notifications through a form on the company’s website.

Category:AI
Did you enjoy this article?
Discover more interesting posts on our blog
Back to blog

Related posts

Altman Announced the Singularity Days After His Models Escaped the Lab on Their OwnAltman Announced the Singularity Days After His Models Escaped the Lab on Their Own
OpenAI chief Sam Altman declared on the Relentless podcast that humanity has already entered the singularity. “We’re like, in the singularity now,” he said verbatim. For decades, the term belonged more to science-fiction literature
6 min read
28. 7. 2026
AI Remixed a Madonna Song—and Now It Tops the Charts in Australia. Musicians Are Furious.AI Remixed a Madonna Song—and Now It Tops the Charts in Australia. Musicians Are Furious.
Since April, Australian radio has been playing a dance remake of Madonna’s hit Like a Prayer on repeat. Released by Queensland DJ Josh Fawaz, it tops the radio airplay chart and has 35 million Spotify streams.
6 min read
28. 7. 2026
Claude Opus 5 Built a Shooter from Scratch. What Can Claude of Duty Do?Claude Opus 5 Built a Shooter from Scratch. What Can Claude of Duty Do?
A first-person shooter that runs directly in the browser, with its own physics and eleven separate code modules. Around 55,000 lines in total, split across eleven subsystems and built on Thr
4 min read
28. 7. 2026
Přihlaste se k odběru našeho newsletteru
Zůstaňte informováni o nejnovějších příspěvcích, exkluzivních nabídkách, a aktualizacích.
CodedTrip

Operated by CodedTrip LLC, USA.

YouTube
TikTok