AI Pulse by Inblix

OpenAI's new model redesigns Nobel-winning proteins with 50x boost

OpenAI Blog · Jul 13, 2026 · 2 min read · Read original article →

Curated by the Inblix editorial team


Featured image for article: OpenAI's new model redesigns Nobel-winning proteins with 50x boost

OpenAI and longevity startup Retro Biosciences just showed what happens when you point a large language model at the protein engineering problem. The result is GPT-4b micro, a custom model they say has redesigned the Yamanaka factors—the Nobel Prize-winning proteins that reprogram adult cells into stem cells—to be dramatically more effective. In lab tests, the AI-generated variants triggered a greater than 50-fold increase in stem cell reprogramming markers compared to the natural versions. That is not a typo.

The Yamanaka factors, discovered by Shinya Yamanaka, are a cornerstone of regenerative medicine, but they have always been frustratingly inefficient. Fewer than 0.1% of cells typically convert, a number that drops even further when you’re working with cells from older or sick patients. The search space for better versions is astronomically large—on the order of 10^1000 possible variants for just two of the four factors. Traditional methods that mutate a few amino acids at a time barely scratch the surface.

This is where GPT-4b micro comes in. The model is a scaled-down version of GPT-4o, fine-tuned on a massive dataset of protein sequences, biological text, and even tokenized 3D structures. Crucially, the training data was enriched with contextual information—evolutionary relationships, known protein interactions, and functional descriptions—that most protein language models ignore. That allows the model to handle intrinsically disordered proteins, like the Yamanaka factors, which don’t have a single stable shape but rely on a flurry of transient interactions to do their job. The team also pushed the context window to 64,000 tokens during inference, a size they call unprecedented for protein models, and found it kept improving controllability.

The redesigned proteins didn’t just boost reprogramming markers; they also showed enhanced DNA damage repair, hinting at a deeper rejuvenation potential. The findings, first observed in early 2025, have now been replicated across multiple donors, cell types, and delivery methods. Retro’s scientists confirmed the resulting stem cells are fully pluripotent and genomically stable. It’s a concrete signal that in silico gains can translate to wet-lab reality, sidestepping the common critique that AI benchmarks for biology don’t mean much in practice.

💡 Key Takeaways

  1. GPT-4b micro redesigned the Yamanaka factors and saw a greater than 50-fold increase in stem cell reprogramming marker expression in vitro, with results replicated across multiple donors.
  2. The model was trained on protein sequences enriched with evolutionary and functional context, allowing it to engineer intrinsically disordered proteins that traditional methods struggle with.
  3. OpenAI observed scaling laws during the model's development, with larger models and datasets yielding predictable performance gains, mirroring how text-based LLMs improve.
  4. The AI-designed proteins also demonstrated enhanced DNA damage repair, suggesting the gains go beyond simple efficiency and may indicate higher rejuvenation potential.

Keep reading: See related articles below for more coverage on this topic.

Get smarter about AI

The sharpest AI news, curated daily. Delivered free to your inbox.

Learn more

Glossary terms

← Back to all articles