GPT-5.6 cracks 30-year stats mystery that stumped humans
Curated by the Inblix editorial team
A 30-year-old assumption in statistics just collapsed in about the time it takes to watch a movie. Edgar Dobriban, a professor at Penn’s Wharton School, used OpenAI’s GPT-5.6 Sol Pro to disprove a long-held belief about one of the most cited tools in science: the Benjamini-Hochberg procedure.
The BH method, introduced in 1995, controls the false discovery rate when researchers run thousands of simultaneous tests — think scanning the genome for disease links. The original paper has racked up over 130,000 citations. For decades, statisticians assumed the method would hold up even when data points were correlated, like genetic variants inherited together. Nobody could prove it. Turns out, they couldn’t prove it because it’s not always true.
Dobriban’s model shows the actual false discovery rate can creep past the target — in one case hitting 0.104 versus the expected 0.1. That’s a small gap, and Dobriban himself says the finding is more theoretical than practical for now. The BH procedure isn’t suddenly useless. But the speed of the breakthrough is what’s rattling the field. GPT-5.6 cracked it in roughly 90 minutes. Its predecessor, GPT-5.5, churned for over 20 hours with multiple agents and got nowhere.
Berkeley statistician Will Fithian called the conjecture “the most interesting open problem in my area of statistics” and framed the result as a marker of AI’s accelerating reach. He also voiced something more personal: “I can’t help but mourn the bygone days when a key result always meant a colleague to celebrate; a human insight to admire; a human achievement to be inspired by.” The model didn’t invent new mathematics — it recombined known methods in an unusual way. That recombination was the hard part, and the newer model found the thread. Whether systems trained on human data can ever produce genuinely new knowledge, rather than remixing what they’ve absorbed, remains the open question looming behind results like this one.
💡 Key Takeaways
- GPT-5.6 Sol Pro disproved a conjecture about the Benjamini-Hochberg procedure in 90 minutes — a problem that resisted human solution for three decades.
- The practical impact is limited for now, but the speed gap between GPT-5.6 and GPT-5.5 (which failed after 20+ hours) signals a real capability jump.
- Berkeley statistician Will Fithian called the disproved conjecture the most interesting open problem in his area, adding that AI-driven results are forcing experts to confront a loss of human-centered discovery.
Keep reading: See related articles below for more coverage on this topic.
Get smarter about AI
The sharpest AI news, curated daily. Delivered free to your inbox.