OpenAI’s Astra is claimed to have solved a math problem
A judge orders the Pentagon to retract its Anthropic bar as Codex tests a mode that keeps working.
By Saeed Ezzati7 min read
The audio edition
Listen to this newsletter
0:003:38
Read transcript
A Cambridge mathematician says OpenAI’s Astra solved the existence problem for non-sofic groups earlier this month. Henry Bradford, a fellow of mathematics at the University of Cambridge, places the claim in group theory—and immediately narrows what it means. In a published letter, Bradford describes the proof as largely a slight twist on earlier theorems by Gabor Kun and Andreas Thom, rather than an entirely new body of theory. That makes this less a clean victory lap than a test of what we call mathematical progress. Bradford says an AI capable of this kind of recombination would have seemed incredible only months ago, and he thinks it would be foolish to bet against superhuman mathematical capability in the coming years. But, drawing on William Thurston’s 1994 writing, he separates producing a theorem from understanding mathematics. The discipline also depends on people learning ideas, sharing them, and extending them. That distinction has an institutional edge. If universities reward high-volume paper production above all else, Bradford worries that cash-strapped administrators could decide human mathematicians are superfluous. He does not present that as inevitable. His argument is that society will have to choose whether mathematics is merely a stockpile of answers, or a human practice of building and transmitting understanding. The Astra claim matters, then, not only because of the result itself, but because it exposes the next evaluation problem: how do we judge machine-generated discovery when the hardest part may be deciding what has actually been understood? That same gap between a promising output and a trusted result appears in medical AI. Researchers in Israel trained a machine-learning model on 97,364 mammograms from 29,921 women, linked to medical records. It reported 86 percent reliability for identifying women who had suffered a stroke, 79 percent for high blood pressure, and 78 percent for coronary heart disease. The potential appeal is practical: one existing screening exam could eventually provide cardiovascular signals without another imaging test. But the study does not define the reliability metric in more detail. Experts say prospective validation, better accuracy, and fewer false results are still needed before clinical use. And when AI expands output, institutions can hit a human bottleneck. London’s National Theatre received 728 applications for one technical apprenticeship; Cox London received more than 160 for two places. Applicants described hands-on craft as less exposed to current generative-AI disruption, though not immune, and sometimes useful alongside AI. The theatre’s Kath Geraghty said many applications now appear AI-generated, making careful review of roughly 700 submissions unviable for staff. The demand shows interest, not completion or employment outcomes. The harder question is whether small apprenticeship programs can turn that interest into durable skills. Finally, MIT’s PottsMPNN pushes in the opposite direction from imitation. Rather than mainly matching protein sequences observed in nature, it evaluates amino-acid choices through structural compatibility and energy prediction. The system models pairwise interactions across all 20 amino acids and adds structural variation during training, aiming to generate sequences that can fold into stable, potentially novel proteins. It still uses evolutionarily related sequences, so the break with nature is not absolute. Across these stories, the practical test is the same: AI can produce impressive candidates, but people and institutions still have to validate them, understand them, and decide what counts as value.


