The debate over artificial intelligence usually splits into two predictable camps: believers waiting for imminent superintelligence and skeptics pointing out that large language models hit a brick wall the moment things get mathematically rigorous. Physicist and science writer Matt von Hippel decided to test that boundary with actual compute rather than philosophical hand-wringing. On his blog, 4gravitons.com, he threw down a very specific gauntlet to AI labs: tackle a calculation that normally eats up months of human specialist bandwidth in theoretical physics.
In theoretical physics, evaluating particle interactions requires computing scattering amplitudes—essentially figuring out how likely subatomic particles are to react based on their momenta and energies. The math is famously brutal, forcing researchers to rely on approximations capped at a specific number of loops representing interaction complexity. Most scattering formulas stall out at two loops, and the gold standard for precision in particle physics sits at five loops. Von Hippel dared the labs to push past that: calculate N=8 supergravity to seven loops, or N=4 super Yang-Mills to nine loops, using standard hardware available to academic researchers.
Computational Frontiers in Scattering Amplitudes
To make these calculations tractable while stress-testing analytical methods, physicists use toy model theories. Yang-Mills theories describe fundamental forces, and in the N=4 supersymmetric version, every particle gets four supersymmetric partners instead of one, creating a clean environment for mathematical torture-tests.
"Give us N=8 supergravity to seven loops, or N=4 super Yang-Mills to nine loops."
Anthropic picked up the challenge and cleared the nine-loop hurdle in the N=4 super Yang-Mills model roughly a month after von Hippel published his post.
What this means: Pulling off a nine-loop calculation in N=4 super Yang-Mills proves that frontier models can manage complex analytical workflows that used to run hard into human compute and time ceilings. For R&D operations, the implications are stark: tasks that previously chewed up months of specialized expert time can now compress down to minutes. However, let us keep the champagne on ice for a moment. This breakthrough happened inside a highly symmetric supersymmetric toy model, not the messy real-world Standard Model. Whether these automated methods can generalize to less cooperative physical systems remains an open question.