Twenty years is five presidential terms, a full academic career, or an eternity in the world of technology. That is exactly how long fundamental statistics lived with a flaw in its foundation—one everyone saw but no one could fix. We are talking about the Benjamini-Hochberg (BH) procedure, the "gold standard" for data processing in genomics and A/B testing. Since 1995, scientists have wondered: does the method remain accurate under complex correlations, when data points "march in formation"? Edgar Dobriban of the Wharton School decided not to waste another two decades on debate and asked GPT-5.6 Pro. The model produced a counterexample, a proof, and the code in ninety minutes. While theorists were defending dissertations based on "compelling evidence," the algorithm simply showed them the door.
The Dobriban Precedent and the Death of Intuition
The mechanics of the process look almost mocking. Dobriban fed the model a mathematical definition and a task: prove or disprove. GPT-5.6 Pro didn’t politely agree with the professor; instead, it constructed a Gaussian factor model using three coordinate blocks. The result: at a set false discovery rate threshold of 0.01, the actual figure was 0.0104. In mathematics, this tiny step over the line signifies a catastrophe—the hypothesis is false. The most important factor here isn't the number, but the "numerical certification": the model provided code that machine-verified the inequality. This is no longer "probabilistic chatter," but cold, high-order reasoning. For comparison: the previous GPT-5.5 tortured the same task for twenty hours with a swarm of agents and produced absolutely nothing.
"A hypothesis that stood for twenty years turned out to be simply wrong. The gap is not an estimate: the final inequality is confirmed by an interval-arithmetic certificate."
For R&D directors, this is a signal: the era of "smart autocorrect" is over. We have entered a phase where AI becomes a lead researcher capable of seeing combinations that the human brain filters out as improbable. As Dobriban himself admitted, the model combined two known techniques that no one had thought to pair before. The problem wasn't a lack of tools, but that humans simply didn't know which way to look.
The Collapse of the Human Factor in Tech-Intensive Business
The real drama unfolded three days after Dobriban’s publication. Lihua Lei of Stanford, who had been struggling with this problem for five years, released his own preprint. In the acknowledgments, he candidly admitted: it was the model's "unexpected counterexample" that resuscitated his research. The productivity dynamics are staggering: a problem stalled for five years was solved in three weeks thanks to a 90-minute "kickstart" from the algorithm. Lei spent 180 hours turning the model's raw insight into ninety pages of proof with precise formulas.
The business impact of such automated intelligence is radical. Imagine testing 20,000 hypotheses simultaneously—searching for markers of a rare disease, for instance. The trusty old Bonferroni correction would stifle the study with impossible requirements, while the Benjamini-Hochberg method promised flexibility but offered no guarantees in real-world, correlated conditions. Now we know: the method fails, and we know exactly by how much. This automated falsification of dogmas eliminates the need for bloated staffs of theoretical analysts whose jobs for years consisted of cautious "ground-testing." Now, a neural network probes the ground for the price of a lunch.
The Limits of Autonomous Science and the Role of Architecture
The qualitative leap of GPT-5.6 Pro represents a transition to hypothesis testing through logical negation rather than text imitation. However, firing the entire R&D department would be premature. The human remains the critical link as the task-setter. Dobriban didn't ask the model to "make it look nice"; he formulated the request in the language of rigorous logic. Without a precise definition at the input, we would have received only another serving of hallucinations. Today, the human is an architect of meaning and a verifier. We have acquired a machine that finishes twenty-year construction projects during a coffee break, but we still need someone to point out which wall needs to come down. The R&D economy is shifting from paying for the search process to paying for the skill of asking the right question.