OpenAI and Los Alamos Put GPT‑4o in a Real Lab to Test Biosecurity Risks
Curated by the Inblix editorial team
Forget theoretical prompts on a screen. OpenAI and Los Alamos National Laboratory (LANL) are taking frontier model testing into the physical world with a new partnership that will see GPT‑4o guide scientists through actual wet lab tasks. This isn’t just another white paper on biosecurity; it’s a first-of-its-kind experiment designed to understand how multimodal AI—models that can see, hear, and speak—might upskill both seasoned researchers and total novices in a real laboratory setting.
The evaluation, led by LANL’s newly formed AI Risks Technical Assessment Group, will task GPT‑4o with assisting in standard protocols like cell culture, genetic transformation, and centrifugation. The goal is to measure exactly how much the model improves task completion and accuracy. As Nick Generous, deputy group leader for Information Systems and Modeling at LANL, put it, the work is about assessing a powerful tool that “,as with any new technology, comes with risks.” The partnership operates under the mandate of a recent White House executive order directing national labs to evaluate frontier AI’s biological capabilities.
This marks a significant and overdue evolution from past safety tests. Previous evaluations, including OpenAI’s own, focused heavily on text-based knowledge—can a model write a convincing synthesis protocol? But knowing the steps in theory is a world apart from pipetting a volatile reagent or visually troubleshooting a contaminated cell culture. GPT‑4o’s voice and vision capabilities change that dynamic entirely. A confused technician can now simply point a camera at a centrifuge and ask why it’s making a strange sound, bypassing the need to translate a messy physical problem into a precise text query.
OpenAI CTO Mira Murati called the partnership a “natural progression” in the company’s mission, and the tie-in with a national lab known for national security work is a clear signal that AI safety is becoming as operational as it is philosophical. The real question this experiment will start to answer isn’t whether GPT‑4o is a biology genius, but whether its real-time sensory input removes so much friction from lab work that it fundamentally lowers the barrier to performing sophisticated—and potentially dangerous—procedures.
💡 Key Takeaways
- This is the first experiment to test a multimodal frontier model's ability to assist with hands-on wet lab techniques, moving beyond text-based theoretical knowledge to real-world benchwork.
- The study will quantify the 'uplift' GPT‑4o provides to both experts and novices, directly measuring how much easier the AI makes it to complete standard biological protocols that act as proxies for dual-use risks.
- GPT‑4o's real-time voice and vision inputs could dramatically simplify troubleshooting, allowing users to get visual guidance on lab equipment without needing to accurately describe the problem in text.
- The partnership is a direct response to a White House executive order tasking national labs with evaluating frontier AI models, blending OpenAI's tech with LANL's national security and biosecurity expertise.
Keep reading: See related articles below for more coverage on this topic.
Get smarter about AI
The sharpest AI news, curated daily. Delivered free to your inbox.