Ex-OpenAI researcher: AI labs will burn $100B on training data because scaling is hitting a wall
Curated by the Inblix editorial team
Andrew Ho spent eight months inside OpenAI. He walked away convinced the industry is chasing the wrong rabbit. The problem isn’t that models need more parameters or compute — it’s that they’re starving for data that actually captures how humans do economically valuable work.
“Most work is highly contextual and not easily encoded into a gradable environment,” Ho writes. He’s not talking about writing sonnets or passing bar exams. He means the messy, tacit knowledge embedded in bioinformatics pipelines, wet lab experiments, and the millions of other domain-specific tasks where current AI fumbles. Even OpenAI’s GPT-5.6 Sol hits roughly a 30 percent success rate on complex scientific analyses, a figure Ho knows firsthand from his time there.
His new startup is a direct bet against the “scale is all you need” religion. The first products target specialized datasets for bioinformatics and routine lab work — think researchers snapping photos of experiments and expecting reliable AI evaluation. Chemistry, materials science, and healthcare are next. Ho pegs the coming spend on targeted data collection at over $100 billion, a number that sounds wild until you look at the economics. Frontier labs like OpenAI and Anthropic, he notes, are chronically unprofitable. They’re forced to sink growing fortunes into each new model generation just to fend off cheaper competitors like Qwen and Kimi.
Ho’s skepticism has company. Cambridge researcher Adam Hunt describes his own trajectory from LLM optimist to pessimist. The latest models, he argues, aren’t becoming more versatile — they’re specializing. Programming and complex math climb while language quality and basic reasoning stagnate or degrade. Google DeepMind’s Tom Zahavy offers a structural diagnosis: language models excel at deduction and induction but can’t perform creative abduction, the leap to invent explanations for which no linguistic precedent exists. His proposed fix involves action-controllable world models that enable counterfactual experimentation. Hunt himself puts his confidence in the scaling-alone thesis at only about 40 percent. The billion-dollar question — whether stitching specialized models together can approximate the general intelligence that pure scale couldn’t deliver — remains completely open.
💡 Key Takeaways
- OpenAI's GPT-5.6 Sol achieves only a 30% success rate on complex bioinformatics tasks, exposing a performance ceiling that more parameters alone won't break through.
- Reinforcement learning thrives in coding because reward signals are clear and data is abundant — most economically valuable skills lack both conditions, which explains why model improvement is wildly uneven across domains.
- Google DeepMind's Tom Zahavy pinpoints creative abduction as the structural blind spot: LLMs can't invent causes that have no existing linguistic precedent, a capability that may require coupling language models with world models that test counterfactuals.
Keep reading: See related articles below for more coverage on this topic.
Get smarter about AI
The sharpest AI news, curated daily. Delivered free to your inbox.