Anthropic recently decided to throw its unreleased research version of Claude into the deep end, tasking it with the Riemann Hypothesis—a 165-year-old mathematical enigma that carries a million-dollar bounty and a reputation for breaking human minds. To the surprise of exactly no one, the model didn’t solve the millennium prize problem. However, the experiment revealed exactly where the current ‘intelligence’ ceiling sits: AI is becoming an elite lab assistant, but it’s still not a pioneer.

Rather than a radical theoretical leap, Claude focused on the grunt work of optimization. By synthesizing decades of research—from Enrico Bombieri’s foundational work to recent papers by Baluyot, Goldston, Suriajaya, and Turnage-Butterbaugh—the model managed to push the lower bound for the fraction of zeros on the critical line from 41.6% to 67.2%. It didn't find a new path; it just built a significantly better bridge using existing blueprints. Anthropic’s internal team, alongside external heavyweights Brian Conrey and Dan Goldston, verified that the proof wasn't just a hallucination but a formally sound mathematical advancement.

This case serves as a reality check for the 'AGI is around the corner' crowd. Claude demonstrated a sophisticated ability to manage complex technical workflows and refine existing methodologies, yet it remains tethered to the data it was trained on. It can iterate on known variables with superhuman speed, but it lacks the creative spark required for a fundamental paradigm shift in 'unsolvable' fields. We are looking at a future where AI handles the heavy lifting of verification and incremental progress, leaving the actual 'eureka' moments to the humans—at least for now.

Large Language ModelsAnthropicGenerative AIArtificial Intelligence