New Era of Voice AI: OpenAI Introduces Advanced Audio Models

New Era of Voice AI: OpenAI Introduces Advanced Audio Models

Ondřej Barták
Ondřej Barták
Entrepreneur and Programmer
24. 3. 2025
4 minutes reading
New Era of Voice AI: OpenAI Introduces Advanced Audio Models

A New Era of Voice AI: OpenAI Introduces Advanced Audio Models

At a time when artificial intelligence is becoming an integral part of our everyday lives, OpenAI has made significant progress in the field of voice technologies. The company recently introduced a new generation of audio models that push the boundaries of what is possible in text-to-speech (TTS) and speech-to-text (STT) conversion. These innovations promise to revolutionize the way we interact with digital assistants and applications.

New Models with Advanced Capabilities

OpenAI has launched three state-of-the-art voice AI models – gpt-4o-transcribe, gpt-4o-mini-transcribe and gpt-4o-mini-tts. These models are designed for high-quality transcription and customizable speech synthesis, opening up new possibilities for both developers and users.

The GPT-4o-mini-tts model represents a major breakthrough in text-to-speech technology. Its key feature is so-called "steerability," which allows developers to control not only the content of a message but also the way it is delivered. Using simple text instructions such as "speak like a pirate" or "use a voice for telling bedtime stories," the model can adapt its speaking style. This feature makes interactions with AI more natural and engaging.

In the field of speech recognition, the GPT-4o-transcribe and GPT-4o-mini-transcribe models achieve unmatched accuracy. With an error rate of just 2.46% for the English language, they surpass existing standards, including OpenAI's previous Whisper models. The models excel particularly in processing different accents, distracting background noises, and varying speech rates.

Multilingual Capabilities and Practical Applications

One of the most significant strengths of the new models is their linguistic versatility. In FLEURS testing, which evaluates transcription accuracy in more than 100 languages, the new models outperformed not only existing Whisper models but also competing solutions. This paves the way for more effectively overcoming language barriers on a global scale.

Unlike Whisper, the new models do not support speaker identification (diarization), but they offer improved noise suppression and semantic voice activity detection. These features are essential for practical real-world applications such as customer support, language learning, and assistive technologies.

OpenAI.fm and Opportunities for Developers

To demonstrate the capabilities of the new models, OpenAI launched the openai.fm platform, where users can try out various AI voice styles in real time. This interactive demo page allows users to experiment with different voice variations and stylizations.

Developers can integrate these models into their applications through the OpenAI API. The company has also improved its Agents SDK, which now makes it possible to transform text-based AI agents into voice agents with minimal coding. This update simplifies the integration of real-time voice interactions into existing applications.

Competitive Landscape and Pricing Policy

With these models, OpenAI is entering a competitive environment that already includes companies such as ElevenLabs with its Scribe product and Hume AI with Octave TTS. Although, according to some evaluations, OpenAI's voices do not yet match the realism of competing solutions such as Sesame or ElevenLabs, their integration into the OpenAI ecosystem represents a significant advantage.

In terms of pricing, OpenAI has set competitive rates: $0.6/min for gpt-4o-transcribe, $0.3/min for gpt-4o-mini-transcribe, and $0.015/min for gpt-4o-mini-tts. For developers using the API, the price is set at $6 per million input audio tokens.

Broader Context and Future Direction

Some critics argue that OpenAI is sidelining real-time conversational AI, while others see this development as an indication of a broader strategy – moving toward full-spectrum multimodal intelligence that would connect AI's text, visual, and audio capabilities into one comprehensive system.

The ability to generate realistic, emotional speech from just a 15-second audio sample, which OpenAI demonstrated through its Voice Engine, suggests where the technology could be heading in the near future.

Why This Matters

OpenAI's latest audio models bring voice interactions with AI closer to natural human conversation, which is essential for their effective use in real-world applications. By enabling greater customizability and expressiveness, these advances help developers create AI agents that communicate more intuitively and can adapt to different user needs.

As these technologies continue to improve, we can expect voice interfaces to become the dominant way of interacting with digital assistants and applications, potentially changing the way we work with technology in everyday life.

At a time when the boundaries between human and artificial communication are becoming increasingly blurred, OpenAI's new audio models represent a significant step toward creating more natural and useful digital assistants that will serve as true helpers in our increasingly complex digital world.

You can try them here and also watch the demonstration video here.

Category:AI
Did you enjoy this article?
Discover more interesting posts on our blog
Back to blog

Related posts

Altman Announced the Singularity Days After His Models Escaped the Lab on Their OwnAltman Announced the Singularity Days After His Models Escaped the Lab on Their Own
OpenAI chief Sam Altman declared on the Relentless podcast that humanity has already entered the singularity. “We’re like, in the singularity now,” he said verbatim. For decades, the term belonged more to science-fiction literature
6 min read
28. 7. 2026
AI Remixed a Madonna Song—and Now It Tops the Charts in Australia. Musicians Are Furious.AI Remixed a Madonna Song—and Now It Tops the Charts in Australia. Musicians Are Furious.
Since April, Australian radio has been playing a dance remake of Madonna’s hit Like a Prayer on repeat. Released by Queensland DJ Josh Fawaz, it tops the radio airplay chart and has 35 million Spotify streams.
6 min read
28. 7. 2026
Claude Opus 5 Built a Shooter from Scratch. What Can Claude of Duty Do?Claude Opus 5 Built a Shooter from Scratch. What Can Claude of Duty Do?
A first-person shooter that runs directly in the browser, with its own physics and eleven separate code modules. Around 55,000 lines in total, split across eleven subsystems and built on Thr
4 min read
28. 7. 2026
Přihlaste se k odběru našeho newsletteru
Zůstaňte informováni o nejnovějších příspěvcích, exkluzivních nabídkách, a aktualizacích.
CodedTrip

Operated by CodedTrip LLC, USA.

YouTube
TikTok