Has it ever happened to you that you interrupted a voice assistant and it completely froze? OpenAI wants to put an end to that. The company is preparing a new voice model codenamed GPT-Bidi-1, and by all indications, it is the biggest improvement to ChatGPT's voice mode in recent months.
The abbreviation refers to a two-way, or bidirectional, architecture that OpenAI has been working on since the beginning of this year. The model is designed to listen and speak at the same time. If you interrupt it mid-sentence, it will not freeze. Instead, it will take your comment into account and adapt smoothly, making the conversation feel more like talking to a person than to a machine that waits until you have finished speaking.
Traces of the new model are now appearing across both the web and mobile versions. That is usually a fairly reliable sign that a rollout to regular users is being prepared. However, the name itself may still change before launch, so take "GPT-Bidi-1" with a grain of salt for now.
OpenAI has allowed a considerable gap to emerge. While its text models have raced ahead all the way to the GPT-5.5 generation, voice has remained on an older audio layer. Spoken conversations have therefore lagged a step behind what the same assistant could do in writing. And the company is not happy about that. It is betting that speech, rather than writing, will become the primary way people interact with artificial intelligence. Its planned audio-focused hardware and voice tools for support are also built around this idea. GPT-Bidi-1 is intended to close this gap. It promises smoother conversations along with a significant leap in reasoning.
What will the feature look like in practice? The outlines are beginning to emerge. ChatGPT users will most likely keep the current setup. They will be able to switch between the new Bidi (Latest) mode and the existing advanced voice mode, so no one will be forced to move to the new system.
The most interesting aspect is the choice of intelligence levels. OpenAI will offer three: high, medium, and instant. This mirrors the tiered system already familiar from the text side and lets people trade speed for depth depending on what they currently need. In a hurry for a quick answer? Choose instant. Want the model to take its time reasoning? Select high.
One recent minor change also suggests that someone is working on a redesign. The voice bubble can now be dragged to the center of the screen. It looks like the first piece of the same puzzle. Exactly when it will arrive remains uncertain. It is impossible to say whether it will launch this week or later. But the groundwork is clearly being laid.
Sources: testingcatalog.com and thewincentral.com



