OpenAI Publishes 722 AI-Generated Math Manuscripts, With Verification Still Uneven
The collection includes computer-checkable proofs and selected reasoning summaries, but OpenAI warns that some unformalized results could contain errors.
Loading page…
The collection includes computer-checkable proofs and selected reasoning summaries, but OpenAI warns that some unformalized results could contain errors.
Listen to this story
OpenAI’s October 6, 2026 release gives researchers a repository to inspect work from its still-unreleased internal mathematics model, but not a uniformly verified corpus: many papers have Lean formalizations, while others do not and may contain problems. The 722 manuscripts are organized into 372 families, and the collection preserves earlier versions as corrections arrive. OpenAI says the work came from an evaluation of about 4,000 open problems; the practical value is access to papers and checkable artifacts, not access to the model or assurance that every result is correct.
OpenAI estimates each average result used compute equivalent to roughly three hours of ChatGPT Pro thinking; this does not give Pro users access to the model.
Only 10 selected results received abridged reasoning summaries, spanning topics from π’s irrationality exponent to a relativistic Vlasov–Maxwell system.
Manuscript families can include companion arguments, consequences, and alternative proofs, so 722 papers do not represent 722 distinct problems.
OpenAI’s internal mathematics model now has a public body of work to inspect: 722 manuscripts, grouped into 372 related families. Published on October 6, 2026, the GitHub collection pairs papers with supporting proof artifacts. But its contents are at different stages of verification, and the model that produced most of them remains unreleased.
Many manuscripts have accompanying formalizations in Lean, a programming language that lets computers check mathematical proofs. Those files provide a different way to examine an argument than reading the paper alone. OpenAI says it will add more formalizations as it obtains them; not every manuscript currently has one.
The repository explicitly warns that some unformalized results could have issues and says OpenAI will endeavor to fix them quickly. The release therefore comes with a verification distinction: it offers manuscripts and, for many, formal proof artifacts, rather than presenting the entire collection as equally checked.
Nor does each paper represent a separate problem solved. A family can collect a principal result, companion arguments, consequences or alternative proofs. Families are classified by mathematical discipline, giving readers a subject-level route into the catalogue before they examine individual manuscripts.
Individual papers in the current catalogue.
Groups of related papers, including companion arguments and alternative proofs.
For readers looking beyond the headline count, the repository supplies several entry points into the papers and their supporting materials:
OpenAI says it expanded its evaluations on open research problems after performance on existing mathematics evaluations saturated. Over the evaluation, the model received approximately 4,000 problems. Outputs were grouped into families and manuscripts, with an appropriate level of significance required for inclusion in the catalogue.
The vast majority of results followed the same procedure using the unreleased internal model. OpenAI estimates that the average result consumed compute equivalent to roughly three hours of ChatGPT Pro thinking. That is the company’s comparison for computing effort, not an announcement that ChatGPT Pro users can access this model.
Alongside the papers, OpenAI released 10 abridged summaries of the model’s reasoning. The selected subjects range from the irrationality exponent of π to a three-dimensional relativistic Vlasov–Maxwell system. These are summaries of selected results, not reasoning accounts accompanying all 722 manuscripts.
Corrections and revisions will appear as new versions, while earlier releases remain accessible. Each manuscript directory supplies a BibTeX citation block. That approach preserves a record of what was published even as papers or supporting artifacts change.
In its release announcement, OpenAI says it consulted the independent Advisory Group on Mathematics and Artificial Intelligence at the Institute for Advanced Study. The company drew on the group’s advice and public recommendations and is exploring community-hosted alternatives for the collection.
OpenAI also promises improvements to citations, mathematical exposition and presentation in future releases. It plans to fund workshops, conferences and special programs to help researchers understand major AI-produced results. Separately, it says it is working toward responsibly releasing the model itself; this publication does not launch it.
Story updates
A major computer science question now has an AI-generated proof claim—and human researchers have already changed their publishing plans under its shadow. OpenAI announced a proof of the Unique Games Conjecture on October 6, 2026, according to Quanta Magazine. The anticipated announcement had pushed three researchers to rush a related result into public view.
The conjecture concerns problems with many rules that must be satisfied at once. Think of coloring a network of points, with each connection imposing a rule about the colors at its ends. Even when a near-perfect solution exists, the conjecture says finding one that satisfies only a tiny fraction of the rules can remain hard. Lowering the standard does not necessarily make the task easy.
Its importance extends beyond that coloring puzzle. In 2008, computer scientist Prasad Raghavendra showed that, if the conjecture is true, one classic algorithm is the best strategy for every constraint satisfaction problem without a perfect solution. Researchers could not improve on it by exploiting a particular problem’s quirks. OpenAI’s announcement included 376 other mathematical results, including 40 other theoretical computer science proofs, Quanta reported.
On September 11, MIT professor Dor Minzer began receiving messages about a rumored OpenAI proof. He and graduate students Yumou Fei and Shuo Wang had not solved Unique Games. They had proved a related result, called the 4-to-1 games result, and were writing it up. Three days after the messages began, they posted a 95-page draft to avoid being overshadowed by OpenAI’s announcement.
Their result addresses a gap in Unique Games: cases where every rule can be satisfied. It is a weaker variant of Subhash Khot’s 2-to-1 games conjecture, allowing four possible colors across a connection rather than two. Yet it carries a long-sought consequence: even for networks that can be colored with three colors, allowing extra colors does not eliminate hard cases where finding a valid coloring remains difficult.
Speed came at the expense of explanation. Minzer told Quanta that, from section six onward, the draft contains definitions and intermediate proofs without connecting prose. The team plans a fuller revision. Their breakthrough had itself emerged from months of unsuccessful attempts: in April 2026, they finally combined lessons from five failures with a new error-correcting code, a method for detecting and fixing errors in transmitted messages.
Math by press release is not that healthy for math
Quanta says OpenAI’s release also included a computer-checked, Lean-verified proof of the stronger 2-to-1 conjecture. That AI-generated manuscript had undergone neither human editing nor independent expert review. Braverman nevertheless saw potential for new research: identifying the proof’s key assumptions could let researchers change them and explore the consequences, though extracting those assumptions from AI-generated work can be difficult.
Minzer’s concern reaches further back in the research process. He argues that failed attempts teach researchers why an approach does not work, and worries that using AI removes that experience. He also fears that the prospect of being overtaken by AI will discourage the ambitious, long-term projects that researchers might otherwise pursue.
quantamagazine.orgLoading discussion...
Join the conversation
Explain whether early scrutiny outweighs the risk of circulating errors.
Be the first to share a perspective or an experience.
Reader comments
Newest comments first. Replies stay oldest first.