OpenAI published 722 mathematical manuscripts on October 6, placing an unusually large body of AI-generated research into a public GitHub repository. The papers are organized into 372 families of related results and were produced by an unreleased internal model that OpenAI is still evaluating.
The release is significant, but the headline numbers need care. A family can contain a main result, companion arguments, consequences or alternative proofs, so 722 manuscripts do not represent 722 distinct discoveries. Nor does publication on GitHub establish that 372 open problems have been solved. OpenAI’s own repository notice warns that the collection contains work at different stages of verification and that some results without formal proofs may contain errors.
What OpenAI has released is closer to a vast research queue: hundreds of claims spanning pure and applied mathematics, supporting source files, a catalog of machine-checkable proofs, and a small set of reasoning summaries. The immediate task now falls to mathematicians and formal-methods researchers who must determine which results are correct, which are important, which duplicate earlier work and which can be explained in a form other researchers can build on.

What is actually in the OpenAI math release
The manuscript catalog ranges across analysis, algebra, geometry, topology, probability, mathematical physics, number theory and theoretical computer science. Examples include claimed bounds for matrix multiplication, work on zero-free regions for the Riemann zeta function, counterexamples to long-standing conjectures, complexity results and a proposed global result for the three-dimensional relativistic Vlasov-Maxwell system.
That range helps explain why a single verdict on the collection would be meaningless. A specialist in operator algebras cannot quickly certify a paper in arithmetic geometry, and a correct proof in one field says nothing about a neighboring manuscript. Each result must be checked against its own literature, definitions and dependencies.
OpenAI reports that the model was presented with about 4,000 problems. After grouping related output and applying what the company describes as an appropriate significance threshold, it retained the 372 families. The average result used compute comparable to about three hours of ChatGPT Pro thinking, according to the company’s disclosure. OpenAI has not released the model itself, however, so outside researchers cannot reproduce the generation process or test how it behaves on a comparable private set.
The repository includes abridged reasoning summaries for only 10 families. Those cover topics including correlations of multiplicative functions, the irrationality exponent of pi, Mahler conjectures, semidefinite-programming hardness, Kaplansky’s direct-finiteness conjecture and the Vlasov-Maxwell system. The summaries offer more process detail than a finished paper alone, but they represent a small fraction of the full catalog.
Why a Lean proof matters, and what it does not prove
Many manuscripts include formalizations written in Lean, a proof assistant that checks whether a theorem follows from explicitly stated definitions and previously accepted results. OpenAI published a machine-readable formalization catalog linking papers to their Lean artifacts and verification configurations.
A successfully checked Lean file is much stronger evidence of logical correctness than fluent mathematical prose. The checker does not become persuaded by a plausible explanation or overlook an algebraic step. If the proof compiles against the declared environment, each formal step satisfies Lean’s rules.
That is not the same as proving that every public claim around the paper is right. Reviewers still need to ask whether the formal theorem accurately captures the informal problem, whether definitions hide assumptions that weaken the result, whether imported lemmas are appropriate and whether the claimed novelty survives a literature search. A formal proof can establish “the encoded theorem follows”; it cannot decide by itself that the encoded theorem is the result readers think they were promised.
The distinction is especially important for the unformalized manuscripts. Traditional expert review remains essential there, and even correct work may need extensive rewriting before other mathematicians can understand the key idea. OpenAI acknowledges that not every paper has a Lean counterpart and promises to add formalizations and record corrections as new versions while preserving earlier releases.
The repository falls short of some requested disclosure standards
The release follows weeks of conflict over credit, research practice and the use of proprietary models on open problems. An independent Advisory Group on Mathematics and Artificial Intelligence, hosted at the Institute for Advanced Study, published recommendations on September 29 after receiving more than 600 responses from the mathematical community.
The group asked labs to publish results promptly, search the literature thoroughly, provide readable exposition, preserve revisions, formalize proofs where possible and disclose the model, prompts, time and estimated compute cost for each result. It also requested an account of failed attempts so readers could judge the system’s success rate rather than see only selected wins.
OpenAI meets parts of that standard. The repository is public, versioning and citation procedures are specified, formal proof artifacts are available for many papers, and the company disclosed the approximate number of attempted problems and average compute. It also promises funding for workshops, conferences and other efforts to help researchers understand major results.
Important gaps remain. The generating model is described only as an internal frontier model. Full prompts and complete traces are not available across the catalog, and the 4,000-problem denominator is too coarse to calculate a meaningful pass rate. The repository is also controlled by OpenAI rather than a community scholarly archive with independent governance and persistent identifiers, though the company says it is exploring alternatives.
The advisory group went further than asking for better documentation: it urged frontier labs to stop testing advanced mathematical problems on proprietary models. Its concern is structural. If only a few companies can run the systems that produce frontier results, they can set the research agenda, consume open problems as evaluations and leave the academic community responsible for verification and explanation.
How to read claims about the 722 papers
Readers should avoid three shortcuts as the collection is reviewed.
- Do not treat manuscript count as breakthrough count. The unit OpenAI uses is a family, and families can contain multiple presentations of related work.
- Do not treat a Lean badge as a complete scientific verdict. Formal checking addresses logical validity of the encoded statement, not novelty, importance, attribution or the match between formal and informal claims.
- Do not treat an error in one paper as a verdict on the collection. The papers cover different fields and use different arguments. Verification will proceed result by result.
The most consequential early signals will come from specialists who can compare particular manuscripts with the existing literature, reproduce the formal builds and explain the central ideas in ordinary mathematical language. Corrections will matter too: a transparent record of changed theorem statements, repaired citations and abandoned claims will reveal more about the quality of the process than the initial paper count.
The model remains unavailable
OpenAI says it is working toward a responsible release of the model behind the papers, but it has not provided a date, access rules or pricing. That leaves the public with the outputs but not the instrument that produced them.
For now, the 722 manuscripts should be treated as claims under review, not a settled ledger of discoveries. The release may contain important advances, and the presence of formal artifacts makes parts of it unusually inspectable for AI-generated research. The harder work begins after generation: establishing correctness, credit, significance and human understanding at a scale mathematics has not previously had to absorb.
Sources: OpenAI, the OpenAI math repository, the Advisory Group on Mathematics and Artificial Intelligence, and The Verge.