Back to Blog

AI Theorem Proving: How Agents Are Changing Math Research

Daily News1838
AI Theorem Proving: How Agents Are Changing Math Research

Introduction

Large language models and AI agent systems have demonstrated remarkable capabilities in advanced mathematical reasoning. A new controversy has emerged within the global mathematics community: AI laboratories are generating valid proofs for long-standing unsolved mathematical conjectures, yet choosing not to publish these results immediately. This practice challenges centuries-old academic conventions that define how mathematical discoveries are shared, credited, reviewed and validated. This article examines the recent wave of revelations, quantitative data on AI proof generation speeds and resource consumption, the four core academic norms under threat, and the philosophical and institutional questions raised by AI as a hidden solver of mathematical puzzles.

The Shock Wave Sweeping the Mathematics Community

Scott Aaronson, a professor of theoretical computer science at the University of Texas at Austin and former visiting researcher with OpenAI’s alignment team, sparked widespread discussion in mid-September. He shared a lighthearted remark from his 13-year-old daughter: if she aims to become a mathematician, she might only have roughly two weeks left before AI dominates mathematical discovery. Aaronson expanded this comment into a long blog post. He compiled a list of long-standing open mathematical problems that AI systems have either proven or helped to partially verify in the preceding month. These problems include counterexamples related to the Babič conjecture, the Fermat’s Last Theorem validated by Lean, and decades-old hard questions within quantum complexity theory.

The more unsettling news came from unconfirmed reports behind closed doors. An AI lab completed a proof for the Navier-Stokes existence and smoothness problem. After facing hostile public reaction to the initial announcement, the institution decided to hold the solution privately while refining its handling strategy. Ethan Mollick, an associate professor at the Wharton School, assessed that the reliability of these reports required careful scrutiny. Meanwhile, Scott Armstrong, a mathematician at New York University, claimed that OpenAI had held hundreds of unpublished mathematical proofs starting as early as the International Congress of Mathematicians (ICM) in July.

This situation represents a dramatic reversal of the historical anxiety in mathematics. For hundreds of years, mathematicians worried that famous hard problems might remain unsolved forever. Today, a new fear spreads through the field: AI may already have derived the answers, but research teams choose to keep these proofs confidential. Hundreds of unpublished proofs sit inside private AI research labs, reshaping the fundamental risk landscape of mathematical research.

AI Turns Theorem Proving into a Pipeline Workflow

Twenty years ago, academic circles celebrated when AI assisted progress on major mathematical conjectures. Now AI can tackle world-class open problems within dozens of hours. On August 1, OpenAI published ten papers covering mathematics and theoretical computer science, announcing solutions or substantial advances on multiple long-standing open questions.

On September 1, internal models built on GPT-6 Astra solved two millennia-old mathematical challenges. A team of nearly 100 AI agents spent about 50 hours generating an unrefined proof. Then the team shifted resources toward the Navier-Stokes problem. Around 10,000 intelligent agents ran concurrently. The preliminary agent cluster took 88 hours to produce the first draft proof, and Lean formal verification consumed an additional 17 hours. This single proof consumed approximately 130 billion output tokens.

Aaronson estimated the computational cost of this 166-page proof at roughly 15 million US dollars in compute resources. To date, few humans have fully read and understood the complete argument. Armstrong noted that modern AI labs can produce a 160-page academic paper alongside Lean formal proofs with hundreds of thousands of lines within just a few days. This level of computational throughput was impossible for top universities in past decades.

The workflow creates a clear separation in timelines. AI can generate a proof within days, but human mathematicians may spend years reading, checking and absorbing the reasoning. Traditional academic publication pipelines cannot match this output cadence. The accumulation of unpublished proofs is a natural consequence of explosive growth in available compute power. This shift undermines four foundational rules that have governed mathematical academia for generations.

Four Established Academic Norms at Risk of Failure

Rule One: Open Publication

Historically, mathematical discoveries become part of public knowledge once formally released. The community can review, debate and build upon the results. Now, Armstrong warns that knowledge transitions from “available for public inspection” to “known only to insiders.” Mathematical exchange, which once relied on open sharing, becomes closed off once AI labs lock breakthroughs behind internal walls. More than one research group possesses unpublished proofs, and OpenAI represents only one prominent example of this trend.

Rule Two: Priority and Credit

The old rule gives credit to the first person who publishes a valid proof. The Navier-Stokes incident broke this convention. A team completed the proof internally, and only contacted mathematicians after they had already generated the result. This reverses the traditional collaborative pattern. Mathematicians typically discuss partial progress openly and exchange ideas as research develops. Post-hoc outreach breaks this long-standing academic tradition.

Rule Three: Pre-registration Before Public Disclosure

Lean formal verification can confirm logical correctness of a proof, but human experts may still struggle to comprehend its reasoning. Aaronson compares journal editors to the gatekeepers described in the novel *The Magicians*. On September 11, 25 Fields Medal winners and other leading scholars signed an open letter. They argued that AI should be treated as an auxiliary tool, not the target of mathematical research. If the cost of publishing results exceeds the value gained by releasing them, labs may simply shelve completed proofs indefinitely.

Rule Four: Authorship and Attribution

The question of authorship grows increasingly tangled. If GPT-6 Astra assisted in creating a proof, should the model be listed in the author block? Mathematicians frequently interact with AI systems during research, yet they struggle to define how to cite or credit the system. If human researchers do not participate in the proof construction process, can they legitimately claim authorship of the final paper? This ambiguity shakes the core definition of mathematical authorship.

When teams run large-scale multi-agent workloads for mathematical reasoning, developers need stable routing and traffic management across model endpoints. 4sapi, an API gateway, helps engineering teams standardize requests and manage access when working with multiple large model services.

Who Holds the Proof, and Who Decides When to Release It?

A PhD student may spend three years struggling on a single hard conjecture. Meanwhile, a complete proof could already sit quietly stored inside the servers of an AI lab, unknown to the student, their supervisor, and peer reviewers.

The old academic contest asks who can prove the theorem first. A new layer is added to this race: who knows which problems have already been solved. When AI agents discover proofs far faster than human researchers, a set of difficult institutional questions emerges. Who decides the release timing? What format will the proof be shared in? Who receives formal credit for the discovery? Who has the authority to confirm whether a proof counts as valid?

Previously, mathematicians worried about the difficulty of solving unsolved problems. Two new anxieties now emerge. Researchers wonder whether the problem they are currently tackling has already been solved by AI in a private lab. They also fear that a partial idea shared casually in discussion may get preempted by AI. Mathematics has historically operated as an open discipline, where ideas are shared freely for collective progress. The field now faces uncertainty over whether that generosity can persist, when complete answers may already be locked away inside private AI infrastructure.

Broader Impacts on Mathematical Research and Academic Institutions

This shift creates a fundamental misalignment between AI production speed and human scholarly workflow. Formal verification systems like Lean can confirm logical soundness, but formal correctness is not equivalent to mathematical understanding. The core value of mathematics lies not only in final answers but also in developing new conceptual frameworks that unlock future discoveries. An AI-generated proof may be logically complete while offering limited new insight for human mathematicians.

The economic incentives for AI labs further complicate the situation. A proof of a famous Millennium Prize problem carries enormous reputational value. Labs may delay publication to prepare accompanying papers, verify edge cases, negotiate intellectual property terms, or avoid hostile public reception. The delay is not caused by the difficulty of proof generation, but by strategic business and risk management decisions.

This creates a two-tier mathematical ecosystem. One tier consists of public mathematicians working in the open, unaware of results hidden in private systems. The second tier is the internal research environment of AI companies, where agents continuously churn out solutions to famous conjectures. This divide risks eroding the incentive structure for human mathematicians. If researchers suspect their target problems may already be solved in secret, the motivation to pursue high-risk long-term problems declines.

The field must develop new social protocols. Possible solutions include independent third-party auditing of AI-generated proofs, time-bound embargo policies, and standardized formal proof repositories. Academic journals and prize committees also need updated rules to handle AI-assisted proofs, define authorship, and establish standards for disclosure. Without updated norms, the core culture of mathematics, built around openness and shared credit, faces gradual erosion.

Conclusion

AI systems have crossed a critical threshold: they can produce complete proofs for some of the most famous unsolved mathematical conjectures. The emerging controversy is not merely about whether AI can prove theorems, but about who controls these results and under what rules they are released. The traditional rules governing credit, publication, authorship and open exchange were built for human-paced research. They struggle to adapt to AI systems that can generate hundreds of proofs within weeks.

This moment forces the mathematics community to reconsider its social contract. Mathematical discovery has always been a collective public enterprise. If answers are held privately within corporate servers, the foundational culture of the discipline changes. It remains to be seen whether new shared standards can emerge to balance the powerful capabilities of AI theorem provers with the centuries-old values of open mathematics. For AI developers building multi-agent reasoning pipelines, these developments also highlight the importance of reliable, controlled access to large model inference endpoints.

International access: https://4sapi.com
Domestic access: https://4sapi.cn

Tags:AI theorem provingAI agentsmathematical reasoningLLM reasoningmulti-agent systems

Recommended reading

Explore more frontier insights and industry know-how.