GPT-4o drops: OpenAI's new omni model reacts as fast as a human
Curated by the Inblix editorial team
OpenAI just pulled back the curtain on GPT-4o, and the ‘o’ stands for omni — a nod to the model’s ability to natively chew on text, audio, and video without breaking a sweat. The headline number? A 232-millisecond response time to audio inputs, averaging 320 milliseconds. That’s not just fast for a machine; it lands squarely in the ballpark of human conversational reaction times. Mira Murati and the team are framing this as a deliberate step toward frictionless human-computer interaction, and on paper, the leap from the old Voice Mode’s glacial 2.8 to 5.4 seconds is frankly absurd.
The real magic trick here isn’t raw speed — it’s architectural. Previous voice interactions were a Rube Goldberg machine: one model transcribes speech to text, GPT-4 or 3.5 processes it, and a third model reads the response aloud. Tone of voice, background noise, laughter? All lost in translation. GPT-4o collapses that entire pipeline into a single end-to-end neural network. That means it can perceive a speaker’s sarcasm, the hum of a coffee shop, or even generate a singing response. OpenAI admits it’s barely begun to explore what this unified model can do, and for once, the cliché about scratching the surface feels earned.
On the benchmark scoreboard, GPT-4o matches GPT-4 Turbo on text and code, with standout gains in non-English languages — a quiet but massive upgrade for global users. The API is also 50% cheaper, continuing the brutal price war that’s making advanced AI feel less like a luxury good. Vision and audio understanding see the most significant jumps, setting new performance records. Yet OpenAI is walking a tightrope with safety. Only text and image inputs are launching publicly; audio outputs will be restricted to preset voices while the company stress-tests the novel risks of a model that can generate emotionally nuanced speech.
Seventy external red-teamers from social psychology, misinformation, and bias disciplines have already tried to break the thing. OpenAI’s Preparedness Framework assessment pegs GPT-4o at no higher than Medium risk across cybersecurity, CBRN threats, and model autonomy. That’s reassuring, but the real stress test starts now, when millions of users start poking at the edges. I’m less worried about the model going rogue and more curious about how quickly developers will build apps that exploit its real-time audio-video chops — and what fresh hell that unleashes for deepfake detection.
💡 Key Takeaways
- GPT-4o's 320ms average audio response time matches human conversation speed, a 10x improvement over the old Voice Mode pipeline.
- Unlike its predecessors, GPT-4o processes text, vision, and audio through a single neural network, preserving nuance like tone and background noise.
- The model matches GPT-4 Turbo on English text and code but significantly outperforms it on non-English languages, with a 50% cheaper API.
- OpenAI is deliberately throttling the audio output launch, limiting it to preset voices while it tackles the unique safety risks of emotionally expressive synthetic speech.
Keep reading: See related articles below for more coverage on this topic.
Get smarter about AI
The sharpest AI news, curated daily. Delivered free to your inbox.