OpenAI Launches Three New Voice AI Models
Curated by the Inblix editorial team
OpenAI just dropped three new voice models that could change how we interact with software. We’re talking GPT-Realtime-2, which has GPT-5-level reasoning and can handle complex requests while keeping conversations natural. Then there’s GPT-Realtime-Translate, a live translator that handles 70+ input languages and spits out 13 output languages in real-time. And finally, GPT-Realtime-Whisper, a streaming speech-to-text tool that transcribes as you speak. The big picture here is that voice is moving beyond simple call-and-response. Developers are building three types of experiences: voice-to-action (like telling an app to find you a house and schedule a tour), systems-to-voice (software proactively telling you about flight delays), and voice-to-voice (AI that goes back and forth with you naturally). Why it matters: This marks a shift from voice as a novelty to voice as a real productivity tool that can listen, reason, translate, and act simultaneously.
💡 Key Takeaways
- GPT-Realtime-2 has GPT-5-class reasoning and can handle complex, multi-step requests conversationally.
- The new models support live translation from 70+ input languages into 13 output languages with no delay.
- Voice AI is evolving into three patterns: voice-to-action, systems-to-voice, and voice-to-voice interactions.
Keep reading: See related articles below for more coverage on this topic.
Get smarter about AI
The sharpest AI news, curated daily. Delivered free to your inbox.