
On August 1, OpenAI pinned an oddly un-product-like list to its research page: an internal build of Astra returned ten results that, in the company’s words, resolve or substantially advance long-standing open problems across high-dimensional geometry, coding theory, group theory, arithmetic circuit complexity, quantum complexity, lattice questions tied to post-quantum cryptography, and extremal combinatorics. The accounting is blunt. Tokens spent finding the solutions would cost roughly $2,000 at Sol API rates. Humans and the same model then turned the arguments into manuscripts; the model formalized each argument into a Lean certificate; and OpenAI is releasing a narration of the model’s thinking for every solution.
This is not soft “the model can do contest math” news. The board has specific nail holes: exponentially improved bounds on the maximum size of binary codes at any prescribed minimum distance, with analogous gains for high-dimensional spherical codes; a construction of non-sofic groups; a disproof of the conjecture that certain groups are uniquely determined by their von Neumann algebras; new lower bounds for computing the permanent, including an arithmetic-formula lower bound of order n⁴/log n; an exponential parallel-repetition theorem for general two-player quantum games; polynomial-factor hardness of approximation for the closest vector problem; a determination, in every dimension, of the maximum volume of a convex body whose centroid is its only interior lattice point (tied to Ehrhart’s volume conjecture); a superexponential lower bound for multicolor triangle Ramsey numbers (Erdős problem 183); and progress on compactness and degeneracy conjectures in extremal graph theory (Erdős problems 146 and 180). Earlier, while evaluating an unreleased model, they had already publicized an AI-generated disproof of the Erdős unit-distance conjecture and said that line of work sparked further developments.
Money and certificates lock the story down. Two thousand dollars is not a cute bargain headline; it reframes frontier mathematical search as a priced inference budget. The scarce resource is less a flash of inspiration on paper than a pipeline that can search, draft, and formalize in a way others can audit. Publishing Lean certificates beside thinking narrations splits “what the model proposed” from “what the machine checked.” The first layer will stay contested; the second can at least be opened as a certificate. OpenAI also draws a hard line on authorship: claiming pure human authorship for a proof generated entirely by an AI system would misrepresent both the system’s contribution and genuine human intellectual work. The company helped prepare manuscripts and formalize proofs, and it takes responsibility for correctness, while the mathematical arguments themselves came from the system—a direct answer to concerns like those around the Leiden declaration on AI and mathematics, not a slogan.
The industry stakes are harder than another leaderboard flex. First, evaluation is sliding from closed benchmarks onto open problems themselves: models are being thrown at unsolved questions during development, and wins feed both the training story and the product story for Astra as the next flagship. Second, formalization becomes a delivery format. Without Lean, claims like these struggle to clear the “are you just talking to yourselves” bar; with certificates, the race becomes who can stably ship inspectable proofs. Third, attribution and priority fights get louder. Journals, conferences, and PhD training that keep pretending the author line is humans-only get hit by results like these; dumping a model into the author list without documenting the pipeline muddies priority just as badly.
The disputes are concrete too. A $2,000 search bill is not the community’s reproduction bill: internal Astra, contamination risk, prompts and toolchains, and compute quotas can turn “ten public problems” into “ten problems only the lab can re-run.” Verified formalization also does not automatically equal mathematical taste or the ability to pose good questions—OpenAI is showing advances and resolutions, not a model that writes the next Erdős list. Place the free ChatGPT seats for 100,000 scientists and mathematicians (ChatGPT for Academic Researchers) next to this drop and the picture sharpens: proofs in one hand, accounts in the other, pulling the mathematical community into both the tool and the authorship politics at once.
My read is plain. What makes this post valuable is not the adjectives in ten titles, but the auditable research delivery chain: search budget → manuscript → Lean → public narration. If later system cards and independent checks hold, part of the bottleneck in mathematical research shifts from “can someone think of it” toward “can someone review it, teach it, and seat a machine proof inside a human theory map.” Reviewers and advisors do not vanish; they look more like quality control and meaning allocation. If you only remember “AI proved ten theorems,” you missed the pipeline OpenAI is actually trying to normalize.