ChatGPT gets eyes and ears, rolling out now to paid users
Curated by the Inblix editorial team
OpenAI is giving ChatGPT the ability to see, hear, and speak, starting with Plus and Enterprise subscribers over the next two weeks. It’s the most significant interface upgrade since the chatbot’s launch, and it fundamentally changes how people will interact with the model — moving from a purely text-based relationship to something that feels closer to a conversation with a person who can also look over your shoulder.
The voice feature, available on iOS and Android, lets users have back-and-forth spoken conversations. Under the hood, it’s powered by a new text-to-speech model that can generate human-like audio from just text and a few seconds of sample speech. OpenAI tapped professional voice actors to create five distinct voices you can choose from, and they’re using their open-source Whisper system to handle speech recognition. The result should feel less like talking to a robot and more like a natural dialogue — request a bedtime story, settle a dinner table argument, or just chat while you’re walking.
On the image side, multimodal GPT-3.5 and GPT-4 can now digest photos, screenshots, and documents that mix text with visuals. Snap your fridge and pantry to get meal ideas with step-by-step recipes. Circle a math problem in a photo and get hints. Point your camera at a broken grill to troubleshoot. There’s even a drawing tool in the mobile app to highlight specific parts of an image you want the model to focus on. The company is positioning this as practical assistance — helping you navigate what you’re literally looking at, not just what you’re typing about.
But the rollout comes with explicit guardrails. OpenAI acknowledges that realistic synthetic voices open the door to impersonation and fraud, which is why they’re limiting the tech to this specific voice chat use case. Spotify is already piloting a similar feature for podcast translation that preserves the host’s voice. On the vision front, the company brought in red teamers to probe for risks around extremism and scientific misuse, and they’ve added technical limits to prevent the model from making direct statements about people in images. They worked with Be My Eyes, a app for blind and low-vision users, to understand when visual analysis is genuinely useful versus when it crosses a line. The underlying message: this is a controlled experiment, not a free-for-all.
💡 Key Takeaways
- Voice and image features are rolling out to paying users first, which gives OpenAI a contained population to stress-test the new capabilities before exposing them to hundreds of millions of free users.
- The text-to-speech model only needs a few seconds of sample audio to generate realistic speech, a capability OpenAI is deliberately restricting to avoid impersonation risks while Spotify uses it for podcast translation.
- OpenAI has hard-coded limitations preventing ChatGPT from analyzing or making direct claims about individual people in images, a privacy measure informed by work with the blind and low-vision community through Be My Eyes.
Keep reading: See related articles below for more coverage on this topic.
Get smarter about AI
The sharpest AI news, curated daily. Delivered free to your inbox.