
On September 4, Anthropic posted a sentence that sounds like science fiction and reads like systems engineering: Claude spent about eleven days, largely autonomously, producing the first end-to-end, computer-checked proof of Fermat’s Last Theorem (FLT). Formalizing Fermat's Last Theorem is concrete on scale—about 13 million lines of Lean, 29,500 intermediate theorems used in the final artifact (roughly 30,300 proved along the way), and about six billion output tokens from an internal research model roughly comparable to Claude Fable 5.1.
Read next to the lab’s recent Riemann-zeta work, the contrast is sharp. There, novelty was new mathematics. Here, novelty is verification—checking a known proof the way a calculator checks arithmetic. Verification is less glamorous than discovery, yet the history of mathematics is full of results that took years to accept, and of theories built on foundations that later cracked.
[1]What the machine finished checking
The official record is clear. The proof follows a simplified Wiles route as exposited by Darmon–Diamond–Taylor. Human mathematical input was mostly occasional high-level nudges from Tianyi Peng—“Jacobian as a scheme sounds high priority,” and the like. The scaffold was Prove2Me, an open collaborative formalization platform from Peng and collaborators at Columbia: a DAG of theorem statements to choose next work, statements separated from proofs to cut Lean compile cost, and natural-language descriptions for search and reuse.
Early attempts failed often. Agents lost global state and stopped collaborating; failed efforts still contributed about 7% of non-boilerplate lines in the final proof. Switching to Prove2Me plus a Claude Code multi-agent harness closed the campaign. Lean used only its three standard axioms; a comparator confirmed the statement matches Mathlib’s FLT statement. Kevin Buzzard, reviewing, said autoformalization artifacts are now “robust enough to be built upon.”
On the same scaffold, three personal Claude Max plans formalized Vinogradov’s Three Primes Theorem in three days—Anthropic’s point that major collaborative formalization need not be limited to research clusters. The full proof is on GitHub.
[1]What formalization actually buys
Human proofs skip steps; Lean does not. Human proofs stand on centuries of literature; formalization stands only on the thin slice already mechanized. Buzzard’s community FLT project was expected to take years; its blueprint alone runs 86 pages. Claude compressed that horizon into roughly two weeks of machine-checkable artifact.
Author’s judgment: what changes is not “can AI do number theory,” but “can refereeing outsource part of the load.” Accepting Perelman, Hales, or weak Goldbach took years. As AI accelerates purported proofs, formalization as default delivery may be the only way the community keeps up. Anthropic itself says a formalized proof should not replace human-readable exposition—but it may be the feasible way to trust AI-scale contributions.
[1]Who benefits, and who is just watching
For mathematicians and formalization communities, this is a stress test of tooling: DAG schedulers like Prove2Me look more like reproducible engineering than a pile of chat tabs. For model labs, Lean loops may also feed discovery—Anthropic says many recent Claude results were formalized in parallel, with partial proofs used to self-check hypotheses. For enterprises and consumers, near-term impact is negligible; this is research-infrastructure news.
Competitively, other labs are also pouring into math assistants. Anthropic’s distinctive move is treating verification as a first-class positive agenda, not only chasing novel theorem leaderboards. Whoever makes “paper + checkable artifact” the default delivery starts defining the credibility layer of post-LLM mathematics.
[1]Columnist view
Reading this as “AI proved Fermat” is wrong. Wiles, Taylor, and the long lineage proved it; Claude compressed a known, notoriously hard-to-check human proof into Lean-acceptable form. Credit belongs to formalization and orchestration, not a rewrite of number-theory history.
Three things matter next. First, the artifact dwarfs Mathlib—autoformalization is still bloated, and community takeover cost is unknown. Second, the 7% failed heritage inside the final repo is a reminder that multi-agent math still advances with scar tissue. Third, when formalization becomes Max-plan affordable, both errors and truths propagate faster—machine checking reduces misses, and can also incentivize dumping formalization ahead of human explanation.
Paired with the Riemann work, the lab looks two-legged: one leg probes open problems; the other welds existing giants into a checkable base. For me, the second leg does more to decide whether mathematics can still trust its own foundations over the next decade.
[1]