
The Riemann hypothesis has stood for more than a century. Anthropic staffer Jarred Sumner — not a mathematician — still told Claude to “take a real stab” at it. In the August 10 research note, Anthropic is clear: the hypothesis itself did not fall. The surprise was a related bound: an unreleased research Claude raised the longstanding lower bound on the share of zeta zeros on the critical line from 41.6% to 67.2%, with write-ups for expert review and a Lean formalization attempt.
[1]The process is spelled out. Round one generated and tried ~650 ideas — all failed. Round two spent about a day and a half coordinating ~60 subagents, ~2,400 shell commands, and hundreds of Python scripts, checking known zeta zeros and refereeing one another. Human input was mostly encouragement (“keep going”). Claude downloaded 54 arXiv papers to check novelty and asked for a human number theorist. Anthropic mathematicians Levent Alpöge and Ralph Furman reviewed; experts Brian Conrey and Dan Goldston also examined the paper on short notice. Technically, Claude combined prior work by Aryan, Baluyot–Goldston–Suriajaya–Turnage-Butterbaugh, and Bombieri to push a new inequality in a function space allowing non-diagonal quadratic forms — across ~31 million output tokens.
[1]Anthropic draws its own line: these techniques are not expected to prove the Riemann hypothesis. What is new is productive misfire — aiming at an impossible target and leaving an auditable advance on a neighboring, checkable claim. Encouraging prompts may have helped the model overcome learned skepticism that AIs should not touch open problems — interesting for products, uneasy for alignment (can external cheerleading systematically change exploration?). Author’s take: value sits in a reviewable bound, not a “AI solved Riemann” headline.
[1]For mathematicians, 41.6%→67.2% becomes a hard case study if it keeps holding under scrutiny. For labs, the subagent split (~2 core idea agents, 13 contributors, 13 validators, 2 draft writers) is a research-orchestration template. Controversies: an unreleased model plus pep-talk prompts is harder to reproduce than open-weight runs, and the public may misread a related bound as solving the hypothesis. Publishing the informal note, expert looks, and an August 13 revised paper is Anthropic trying to cool that misread.
[1]Next to later FLT formalization, this episode looks like walking into the wrong door and finding a side window open. Score math agents less on holy-grail hits and more on whether neighboring, checkable propositions gain increments experts will sign.
[1]