A New Era of Voice AI: OpenAI Introduces Advanced Audio Models
At a time when artificial intelligence is becoming an integral part of our everyday lives, OpenAI has made significant progress in the field of voice technologies. The company recently introduced a new generation of audio models that push the boundaries of what is possible in text-to-speech (TTS) and speech-to-text (STT) conversion. These innovations promise to revolutionize the way we interact with digital assistants and applications.
New Models with Advanced Capabilities
OpenAI has launched three state-of-the-art voice AI models – gpt-4o-transcribe, gpt-4o-mini-transcribe and gpt-4o-mini-tts. These models are designed for high-quality transcription and customizable speech synthesis, opening up new possibilities for both developers and users.
The GPT-4o-mini-tts model represents a major breakthrough in text-to-speech technology. Its key feature is so-called "steerability," which allows developers to control not only the content of a message but also the way it is delivered. Using simple text instructions such as "speak like a pirate" or "use a voice for telling bedtime stories," the model can adapt its speaking style. This feature makes interactions with AI more natural and engaging.
In the field of speech recognition, the GPT-4o-transcribe and GPT-4o-mini-transcribe models achieve unmatched accuracy. With an error rate of just 2.46% for the English language, they surpass existing standards, including OpenAI's previous Whisper models. The models excel particularly in processing different accents, distracting background noises, and varying speech rates.
Multilingual Capabilities and Practical Applications
One of the most significant strengths of the new models is their linguistic versatility. In FLEURS testing, which evaluates transcription accuracy in more than 100 languages, the new models outperformed not only existing Whisper models but also competing solutions. This paves the way for more effectively overcoming language barriers on a global scale.
Unlike Whisper, the new models do not support speaker identification (diarization), but they offer improved noise suppression and semantic voice activity detection. These features are essential for practical real-world applications such as customer support, language learning, and assistive technologies.
OpenAI.fm and Opportunities for Developers
To demonstrate the capabilities of the new models, OpenAI launched the openai.fm platform, where users can try out various AI voice styles in real time. This interactive demo page allows users to experiment with different voice variations and stylizations.
Developers can integrate these models into their applications through the OpenAI API. The company has also improved its Agents SDK, which now makes it possible to transform text-based AI agents into voice agents with minimal coding. This update simplifies the integration of real-time voice interactions into existing applications.
Competitive Landscape and Pricing Policy
With these models, OpenAI is entering a competitive environment that already includes companies such as ElevenLabs with its Scribe product and Hume AI with Octave TTS. Although, according to some evaluations, OpenAI's voices do not yet match the realism of competing solutions such as Sesame or ElevenLabs, their integration into the OpenAI ecosystem represents a significant advantage.
In terms of pricing, OpenAI has set competitive rates: $0.6/min for gpt-4o-transcribe, $0.3/min for gpt-4o-mini-transcribe, and $0.015/min for gpt-4o-mini-tts. For developers using the API, the price is set at $6 per million input audio tokens.
Broader Context and Future Direction
Some critics argue that OpenAI is sidelining real-time conversational AI, while others see this development as an indication of a broader strategy – moving toward full-spectrum multimodal intelligence that would connect AI's text, visual, and audio capabilities into one comprehensive system.
The ability to generate realistic, emotional speech from just a 15-second audio sample, which OpenAI demonstrated through its Voice Engine, suggests where the technology could be heading in the near future.
Why This Matters
OpenAI's latest audio models bring voice interactions with AI closer to natural human conversation, which is essential for their effective use in real-world applications. By enabling greater customizability and expressiveness, these advances help developers create AI agents that communicate more intuitively and can adapt to different user needs.
As these technologies continue to improve, we can expect voice interfaces to become the dominant way of interacting with digital assistants and applications, potentially changing the way we work with technology in everyday life.
At a time when the boundaries between human and artificial communication are becoming increasingly blurred, OpenAI's new audio models represent a significant step toward creating more natural and useful digital assistants that will serve as true helpers in our increasingly complex digital world.
You can try them here and also watch the demonstration video here.



