AI Pulse by Inblix

GPT-4o's Reasoning Plummets 26 Points When Listening Instead of Reading

Hugging Face Blog · Dec 20, 2024 · 2 min read · Read original article →

Curated by the Inblix editorial team


Featured image for article: GPT-4o's Reasoning Plummets 26 Points When Listening Instead of Reading

Give a state-of-the-art model a logic puzzle in writing, and it aces it. Read the same puzzle aloud, and that performance collapses. That’s the stark finding from Artificial Analysis’s new Big Bench Audio benchmark, which reveals a massive “speech reasoning gap” in frontier AI models.

OpenAI’s GPT-4o scored a near-perfect 92% on a text version of the test, which adapts 1,000 questions from the rigorous Big Bench Hard dataset into audio. But when the model had to listen to the question and speak its answer directly—a native speech-to-speech interaction—its accuracy plunged to just 66%. That’s not a minor dip; it’s a 26-percentage-point failure that suggests these audio interfaces are fundamentally less reliable for complex thinking.

The experiment, which analyzed GPT-4o and Google’s Gemini 1.5 series, tested four configurations: speech-to-speech, speech-to-text, text-to-speech, and text-to-text. The questions spanned deductive logic, spatial navigation, object counting, and boolean reasoning—all generated with 23 synthetic voices. To grade the responses, researchers used Anthropic’s Claude 3.5 Sonnet as an automated evaluator, checking for consistency with a ground-truth answer.

While the text-only baseline shows these models can reason, the audio modality clearly introduces a bottleneck. The study doesn’t just flag a weakness—it exposes a design reality where the convenience of voice interaction may come at a steep cognitive cost. For anyone planning to talk through a serious problem with an AI assistant, the results are a warning: you might be better off typing it out.

💡 Key Takeaways

  1. GPT-4o's accuracy on a hard reasoning test drops from 92% in text-to-text mode to 66% in native speech-to-speech mode, a 26-point gap.
  2. The Big Bench Audio benchmark uses 1,000 audio questions adapted from Big Bench Hard, spanning formal logic, navigation, counting, and boolean reasoning.
  3. The study used Anthropic's Claude 3.5 Sonnet to automatically grade model responses, establishing a scalable method for audio benchmarking.

Keep reading: See related articles below for more coverage on this topic.

Get smarter about AI

The sharpest AI news, curated daily. Delivered free to your inbox.

Learn more

Glossary terms

← Back to all articles