AI Pulse by Inblix

OpenAI's worst-case fine-tuning test found GPT-OSS still trails o3 on dangerous capabilities

OpenAI Blog · Jul 13, 2026 · 2 min read · Read original article →

Curated by the Inblix editorial team


Featured image for article: OpenAI's worst-case fine-tuning test found GPT-OSS still trails o3 on dangerous capabilities

Before OpenAI released GPT-OSS, they tried to break it. Not casually — systematically. The company’s new paper details ‘malicious fine-tuning’ (MFT), a red-team methodology designed to squeeze every last drop of dangerous capability out of an open-weight model before deciding whether it’s safe to release. The question wasn’t whether the model could be misused, but how bad the misuse could possibly get.

The researchers trained GPT-OSS in two nightmare scenarios. For biological risks, they curated threat-creation tasks and let the model loose in an RL environment with web browsing — essentially giving it tools and incentive to figure out how to cause harm. For cybersecurity, they dropped it into an agentic coding setup to solve capture-the-flag challenges, the kind of hands-on hacking puzzles that measure real offensive capability. These aren’t theoretical benchmarks. They’re active attempts to build a worst-case version of the model.

The results are genuinely reassuring, and not in the vague press-release way. The MFT-trained GPT-OSS consistently underperformed OpenAI’s own o3 model, which the company already rates below its Preparedness Framework’s ‘High’ capability threshold for both biorisk and cybersecurity. Against other open-weight models, the fine-tuned GPT-OSS showed marginal gains in biological capabilities, but nothing that moves the frontier in a meaningful way. ‘Does not substantially advance the frontier’ is the paper’s clinical phrasing — and for anyone watching the open-weight debate, that’s the sentence that matters.

This is the kind of work that should accompany every major open-weight release, and almost never does. OpenAI is publishing the MFT methodology in hopes it becomes a standard part of the release evaluation toolkit. The subtext is clear: if you can’t demonstrate you’ve tried and failed to weaponize your own model, you haven’t done enough homework to put it in the wild. Whether other labs pick up that gauntlet is an open question, but the framework now exists with a name and a paper behind it.

💡 Key Takeaways

  1. OpenAI's malicious fine-tuning methodology actively trained GPT-OSS to maximize harm in biology and cybersecurity, yet the resulting model still fell short of o3 — which OpenAI itself doesn't consider dangerously capable.
  2. The paper introduces MFT as a repeatable evaluation framework for open-weight releases, shifting the burden of proof from 'is this safe' to 'we tried to make it dangerous and couldn't.'
  3. GPT-OSS showed only marginal gains in biological capability over existing open-weight models, with the paper explicitly stating it 'does not substantially advance the frontier' for either risk domain.
  4. The research directly influenced OpenAI's release decision, signaling a move toward evidence-based risk assessment rather than precautionary withholding of open-weight models.

Keep reading: See related articles below for more coverage on this topic.

Get smarter about AI

The sharpest AI news, curated daily. Delivered free to your inbox.

Learn more

Glossary terms

← Back to all articles