In September 2026, the math world learned a question it had never asked before
Not "who will prove it?" but "has it already been proven — and are we allowed to know?"
On September 15, Scott Aaronson — a theoretical computer science professor at UT Austin who once served as a visiting researcher on OpenAI's alignment team — published a blog post containing a joke from his 13-year-old daughter: "If I wanted to be a mathematician, it looks like I have about two weeks left."
It reads like a punchline, but the post was not joking. Aaronson listed open problems proved or assisted by AI over the previous month: a counterexample related to the Jacobian conjecture, a Lean-verified formalization of Fermat's Last Theorem, long-dangling problems in quantum complexity theory. The item that unsettled mathematicians came at the end of the list.
According to Aaronson, after the hostile public reaction to the Navier-Stokes proof, AI labs are now sitting on solutions to "very major problems" until they figure out how to handle them. Scott Armstrong, a mathematics professor at New York University, put a number on it: he was told that OpenAI has been sitting on "hundreds" of proofs since at least the International Congress of Mathematicians (ICM) in July.
The popular read is "AI wins again." The real story is structural: for the first time, the answer is outrunning human understanding — and the answers are locked inside private servers. That is not gossip. That is a crack in the machinery of knowledge itself.
Machines measure output in days. Humans measure comprehension in years.
Start with the scale of the mismatch. On August 1, OpenAI released ten mathematics and theoretical computer science results, each claiming to resolve or substantially advance a long-standing open problem. On September 1, OpenAI heard that two Millennium Prize Problems had been cracked by someone else. Its response was blunt: feed every unsolved Millennium Problem, plus a batch of high-impact questions, to an internal model stronger than GPT-6 Astra.
A group of roughly 100 agents dispatched the force-free Euler equations in about 50 hours. Resources then concentrated on Navier-Stokes: about 10,000 concurrent agents ran for 88 hours to produce a proof, followed by 17 hours of Lean formal verification. That single problem consumed roughly 130 billion output tokens. Aaronson estimates the 166-page proof burned at least $15 million in compute — and almost no human has actually read it.
Now compare the other side of the ledger. Under the Clay Mathematics Institute's rules, a Millennium Prize requires formal publication, a two-year waiting period, and intense external review. As of this writing, Clay still lists Navier-Stokes as unsolved. Machines count their output in days; humans count theirs in years.
Armstrong reports that labs can now produce a 160-page paper plus a Lean formalization running to hundreds of thousands of lines within days. Peer review, conference talks, textbooks — a pipeline refined over centuries simply cannot absorb that cadence. A backlog of unpublished proofs inside AI labs is not scandal; it is the arithmetic consequence of an compute explosion meeting a fixed-speed review system.
Four old rules, broken at once
Academic science runs on four implicit rules: priority goes to whoever publishes first, credit goes to whoever does the work, results must be public, and peers must vet them before they count. The Navier-Stokes affair cracked all four simultaneously.
First, publicity. Knowledge is shifting from "anyone can look it up" to "insiders hear it first." Armstrong writes that rumors about major results fly everywhere, and how much you know depends on your social distance from the lab. Mathematics advanced for three centuries on open exchange. Locking breakthroughs behind lab doors is a regression, not an acceleration.
Second, priority. Tristan Buckmaster (NYU) and Levent Alpöge (Anthropic) had already made progress in the relevant direction. OpenAI heard the rumors, overwhelmed the problem with compute in under a week, and only then contacted the two mathematicians to propose a joint release. Terence Tao warned on Mastodon that merely letting slip that you are "working on a problem" can now invite a compute-driven invasion that flattens the human project before it matures. The likely result: researchers stop sharing promising ideas with colleagues — reversing centuries of open science.
Third, read-before-publish. Lean guarantees the logic of a proof, not that any human understands it. Aaronson compares journal editors to the defenders of Minas Tirith in The Lord of the Rings: without delegating some reviewing to AI, they simply cannot hold the wall. On September 11, Terence Tao and 25 Fields Medalists published an open letter accusing AI companies of treating open math problems as benchmark fodder. Its sharpest line: the goal of mathematics is understanding; solving is merely the means. Here is the bind — publish fast and nobody understands, publish slowly and you are accused of hoarding. The cost of releasing an answer has, for the first time, exceeded the cost of finding it.
Fourth, authorship. If you contributed nothing to the search for the proof, can your name go on it? What does it mean to list GPT-6 Astra as an author? Nobody has answers yet.
A framework: the Proof Gap
Call the reusable tool here the Proof Gap: the widening distance between the speed at which machines produce answers and the speed at which humans verify, understand, and institutionally digest them. It has three layers:
- Production layer: the cost of generating answers has collapsed. A Millennium Problem went from "generations" to "88 hours and 130 billion tokens."
- Verification layer: the cost of confirming answers is rising. A 166-page proof plus hundreds of thousands of lines of Lean code, read by no human, makes Clay's two-year review window almost decorative.
- Institutional layer: the rules for credit and dissemination fail. Priority, authorship, publicity, and peer review all assumed scarce answers visible to everyone.
The danger is not the gap itself but the absence of any synchronization mechanism across the layers. When answers sit on private servers, the institutional layer degrades: knowledge turns from a public good into private inventory, and proximity to a lab becomes the new academic privilege. Read Aaronson's line carefully: the singularity has begun, but it is unevenly distributed. Uneven not just in problem-solving power, but in access to information — and in the very sense that doing mathematics matters.
Noam Brown, an OpenAI researcher, confirmed the pattern from the inside on a recent Dwarkesh Podcast interview: OpenAI's internal models can already answer multiple previously unsolved math problems that the outside world cannot access. "This is a situation with an unfair advantage," he admitted, "and I don't know how to weigh the trade-offs properly." He went further: as progress accelerates, labs may decide internal use is valuable enough to stop external deployment entirely. "By default, AI's external deployment will lag its internal deployment significantly in quality."
The gap is not just about math
If you read this as an internal crisis of mathematics, you are underestimating its portability. Any field where answers are verifiable quickly but understanding takes time will replay the same scissors.
In drug discovery, AI screens candidate molecules in days, but clinical validation takes a decade — the gap exists today, hidden behind regulatory process. In software security, AI finds critical vulnerabilities in days while patching and ecosystem coordination take months — disclosure policy faces the exact dilemma mathematicians now face: publish and risk exploitation, withhold and risk permanent hoarding. In law, AI already organizes case law faster than judges and lawyers can absorb new precedents.
The shared structure: AI's impact is never "who is smarter." It is the time constant between production and digestion in each field. The more a field's digestive process depends on slow human consensus, the sooner its institutions tear.
What to do about it
- If you are a researcher: re-evaluate your "grind" problems — a PhD student who spent three years on a problem may find the answer already sitting on a lab server. Prioritize what the production layer cannot reach: posing new questions, building new subfields, experimental science. And put Lean-style formal verification into your workflow; a proof that machines cannot verify will soon not count.
- If you run an institution (university, journal, funder): peer review must start absorbing AI assistance, and disclosure pressure must be built around private knowledge inventories. Write the Fields Medalists' principle into policy: fund and promote understandable knowledge, not merely claimable answers.
- If you are a practitioner (software, law, medicine): identify your field's digestion bottleneck — audit, review, institutional design. Those "verification-side" roles become scarcer and more valuable as the gap widens.
- If you are an ordinary reader: when you see "AI solved X," ask one question: has this answer been independently verified, or does it live only on one company's servers?
Academia used to compete on who proved it first. From now on, it also competes on who knows what has already been proved. Mathematics has been the most generous of human disciplines — one proof written down, the whole world can read it and build on it. Whether that generosity survives depends on how fast the institutional layer closes the Proof Gap. The window is likely narrower than 88 hours.
