
For two weeks, OpenAI had claimed that an internal model solved more than 100 open problems in mathematics without showing the proofs behind the claim. On October 6, the company delivered: 722 manuscripts collected in a public GitHub archive, some of them with machine-checked proofs. For the first time, the field can see for itself what is behind the announcements. At the same time, the release shows how poorly mathematics is prepared for research at this pace.
Key takeaways
- The archive contains 722 manuscripts in 372 thematic families spanning pure mathematics, theoretical computer science and mathematical physics.
- OpenAI posed roughly 4,000 problems to the unreleased model. Each result took about three hours of ChatGPT Pro thinking compute on average.
- According to the catalog, 162 manuscripts come with a formalization of their main result in Lean, a proof language a computer can check. The rest do not; OpenAI concedes that some results could contain errors.
- An independent advisory group of mathematicians had recommended rules for the release. OpenAI followed them only in part.
What is in the archive
The archive is plainly organized. An overview document describes the 372 families, meaning groups of related papers with a principal result, consequences or alternative proofs. Each paper comes as a PDF with source files and citation instructions. For ten results, OpenAI also publishes abridged summaries of the model’s reasoning. The titles of those families show how high the company is aiming: they range from the irrationality exponent of pi and the conjectures of mathematician Kurt Mahler to the isomorphism problem for free group factors, a question in operator algebras that has been open for decades. They also include a counterexample to a conjecture by the algebraist Irving Kaplansky, whose core is formalized in Lean.
OpenAI describes only briefly how the results came about. The vast majority came from a fixed procedure using the same internal model. After the model saturated the company’s existing math evaluations, OpenAI expanded testing to open research problems. Of the roughly 4,000 problems posed, just under a tenth remained after grouping and applying a minimum bar for significance. The company itself names two prominent exceptions: a paper on a zero-free region of the Riemann zeta function, whose text was edited by humans for readability, and a paper on the Hodge conjecture for a special class of geometric objects. Neither followed the standard procedure.
Lean makes the difference, but only in part
The real advance over earlier announcements is the Lean proofs. Lean is a programming language in which mathematical proofs can be written so that a computer checks every single step. If a result is cleanly formalized there, no one has to take the model’s word for it. Experts then only need to verify that the formal statement really matches what is being claimed. The archive provides such a formalization for 162 manuscripts, and OpenAI plans to add more.
That does not apply to the remaining 560 or so manuscripts. For those, only traditional human review remains, and it is slow. A single demanding proof can keep experts busy for weeks. OpenAI itself writes that some of the unformalized results could have issues and says corrections will be recorded as new versions without deleting older ones. Citations and exposition also still need work, the company says. On top of that, the model that produced the papers is not public. Independent researchers can check the results but cannot reproduce how they were generated.
An advisory group with recommendations but no power
As we reported after the announcement of more than 100 solved problems, OpenAI set up an advisory group at the Institute for Advanced Study in Princeton, the Advisory Group on Mathematics and Artificial Intelligence. Its nine members are unpaid and have no decision-making authority. On September 29, the group published guidelines based on more than 600 responses from the field. According to reports on the guidelines, they include depositing results in an academic repository not controlled by any AI lab and disclosing, for every result, the model name, prompts, reasoning, time spent and compute cost.
OpenAI implemented some of this and not the rest. The archive sits on the company’s own GitHub account, although OpenAI says it is exploring community-hosted alternatives. Prompts for individual problems are missing, and reasoning summaries exist for only ten results. In its statement of October 6, the advisory group put it diplomatically: the release marks the beginning, not the completion, of the process of human understanding. Whether its recommendations were followed successfully, it said, is ultimately for the mathematical community to judge.
Frustration in parts of the field goes back further. With the purported Navier-Stokes proof in September, the dispute was already about attribution and how human contributions were handled. Now concerns about capacity are added: who is supposed to review more than 700 papers when the next batch may already be in the works? And who is responsible for an error when the only author listed is a company?
What this means for mathematics
For researchers, the archive is first of all a toolbox. Anyone working on one of the topics can use the manuscripts as a template, as a counterpoint or as a starting point, and the formalized results can be pulled into their own Lean projects. The formalizations are licensed under Apache 2.0, which makes them open. In subfields with few specialists, that could speed up the work.
The bigger shift, however, concerns the role of humans. Until now, a theorem counted as proven when experts had understood and accepted the proof. With hundreds of machine-generated papers arriving at once, that balance tips: understanding lags behind production. Formal verification in Lean solves part of this problem, because a checked proof is correct even if no one has followed it yet. But mathematics does not live on correct theorems alone; it lives on ideas that others can build on. Whether the 722 manuscripts contain such ideas can only be found out the slow way. OpenAI has announced that it will fund workshops and conferences on the most important results. How serious that commitment is will be measured by how much time and money the company puts into exactly that slow part.
Sources
- OpenAI auf GitHub: openai/math (Manuskripte, Lean-Formalisierungen, README)
- heise online: OpenAI – Hunderte weitere KI-generierte Lösungen für Mathematikprobleme
- Advisory Group on Mathematics and Artificial Intelligence: Empfehlungen und Stellungnahme
- TokenPost: OpenAI Releases 722 AI-Generated Math Manuscripts After Criticism
- Runtime Wire: OpenAI publishes 722 math manuscripts generated by an unreleased model
- Interesting Engineering: OpenAI's largest math release tackles 4,000 problems with Lean proofs

