OpenAI releases 722 AI-written math papers
Many include computer-checkable proofs, but verification is uneven and the model itself remains unreleased.
By Saeed Ezzati8 min read
The audio edition
Listen to this newsletter
0:003:29
Read transcript
OpenAI has published a catalogue of 722 mathematics manuscripts, grouped into 372 related families, produced largely by an internal model the company has not released. Researchers can inspect the papers and supporting files; they cannot run the model. And 722 is a paper count, not a count of distinct problems solved: a family may include a main result, companion arguments, consequences, or alternative proofs. The release is also unevenly verified. Many papers include formalizations in Lean, a programming language that lets computers check proofs. Others do not. OpenAI warns that some unformalized results could have issues, and says it will add formalizations and publish corrections as new versions while keeping earlier releases available. So this is a research record with explicit caveats, not a claim that every result has been independently confirmed. OpenAI says its existing math evaluations had saturated, so it expanded testing to open research problems. The model received about 4,000 problems during evaluation; only work meeting an appropriate significance bar entered the catalogue. The repository includes PDFs, source files, build instructions, and verification configurations where available. The company also published abridged reasoning summaries for just 10 selected results—not all 722. OpenAI says it consulted an independent mathematics advisory group, plans workshops and better citations, and is working toward a responsible model release. For now, the useful next step is scrutiny of the artifacts: especially which claims gain checkable proofs, and which need correction. That distinction between a promising tool and a verified result matters in public health, too. Google says World Health Organization teams in the Democratic Republic of Congo used a conversational mapping prototype to identify 48 settlements and more than 45,500 people considered at risk of Ebola. The mapping reportedly took minutes rather than weeks, helping teams plan mobile-lab deployment and border surveillance. Separate models estimated spread into uninfected areas. Google describes these emergency tools as research prototypes, distinct from its commercially available population dataset. On the security side, Anthropic is widening access to Claude through three tiers for vetted teams: defensive work, authorized penetration testing, and limited testing of safety-critical systems. In a 50-trial-per-tier evaluation, Defense Access blocked 46 trials; Red Team Access blocked none and completed 34. Those are benchmark results, not a measure of real-world misuse. Access still requires verification and monitoring, and Anthropic says blocks remain for threats such as ransomware and disruptive physical-system testing. The practical question is whether broader permission can support legitimate testing without weakening those safeguards. And in clinical care, NYU Langone researchers tested an AI model combining heart recordings with two blood-test biomarkers for transplant rejection. In a 38-person test group, it correctly identified 94 percent of patients without rejection; that is not an overall accuracy score. Researchers projected that the combined model could have spared 19 patients flagged by a blood-test model from biopsies, but no procedures were reported as actually avoided. They say the next step is testing more patients across multiple transplant centers. Across these stories, watch what happens after access: whether mathematical claims gain checkable proofs, prototypes hold up in real workflows, security safeguards work, and clinical findings replicate.




