OpenAI's GPT-4o safety report: Persuasion gets a 'medium' risk rating
Curated by the Inblix editorial team
OpenAI just dropped the System Card for GPT-4o, its new omni model that handles text, audio, and images natively. The headline from their internal safety scorecard? Most frontier risks landed squarely in the “Low” category, but one capability — persuasion — skated right up to the edge, scoring a borderline “Medium.” That’s the one keeping safety teams up at night, not some Skynet autonomy scenario. Cybersecurity, biological threats, and model autonomy all rated Low.
The company is framing this release as part of its voluntary White House commitments, publishing both the Preparedness Framework scorecard and the full System Card to offer what they call an “end-to-end safety assessment.” The real novelty here isn’t the text or vision evals — those build on GPT-4 and GPT-4V work — but the audio capabilities. Voice introduces genuinely new risks around speaker identification, unauthorized voice generation, and ungrounded inference. Their findings suggest the voice modality doesn’t meaningfully amplify Preparedness risks, but they’ve layered safeguards at both the model and system level anyway.
Technically, GPT-4o is an autoregressive model that processes everything through a single neural network, hitting response times as low as 232 milliseconds — roughly human conversational speed. It matches GPT-4 Turbo on English text and code, improves dramatically on non-English languages, and costs 50% less via the API. The training data runs through October 2023 and pulls from web crawls, proprietary partnerships like Shutterstock, and multimodal sources including video.
What’s notable is OpenAI’s candid admission that pre-training data filtering alone can’t solve safety. The bulk of effective mitigation happens post-training, through human preference alignment, red teaming, and product-level monitoring. The Safety Advisory Group reviewed the evaluations before signing off on deployment. No critical risks flagged, but that medium persuasion score suggests the model’s ability to sway opinions or shape beliefs is measurably stronger than other frontier capabilities — and that’s exactly where the regulatory conversation has been heading.
💡 Key Takeaways
- Persuasion was the only Preparedness Framework category to approach Medium risk, while cybersecurity, biological threats, and model autonomy all scored Low.
- GPT-4o's voice modality introduces novel risks like speaker identification and unauthorized voice generation, but internal testing found it doesn't materially increase frontier risk levels.
- The model processes audio, text, and images through a single neural network with response times as low as 232 milliseconds — matching human conversation speed — and costs 50% less than GPT-4 Turbo.
- OpenAI acknowledged that pre-training data filtering is insufficient for safety, with the majority of effective mitigations happening during post-training alignment and red teaming.
Keep reading: See related articles below for more coverage on this topic.
Get smarter about AI
The sharpest AI news, curated daily. Delivered free to your inbox.