DeepMind agents got fake math proofs accepted in 27 minutes
The simulation also produced whistleblowers, but they lacked the power to remove fraudulent submissions.
By Saeed Ezzati8 min read
The audio edition
Listen to this newsletter
0:003:49
Read transcript
The clearest warning today comes from Google DeepMind’s 100-agent math experiment: a system can enforce the appearance of rigor while losing the substance. In a simulated scientific conference, 100 Gemini 3.1 Pro agents worked on 71 formal conjectures in Lean, with access to a public forum, direct messages, and a shared library. They honestly solved 37. Then one agent found a notation-shadowing exploit in the Lean 4 grader. The evaluator checked whether submitted code compiled and looked formally correct, but not whether it actually proved the original claim. Once that exploit entered the shared workspace, the other 34 problems were accepted with fabricated proofs in 27 minutes. Accepted work was automatically added to the library, so the grading flaw became shared infrastructure. Other agents could reproduce it, while honest agents lost access to problems that were now marked solved. The response was mixed. Some agents cheated, some kept working honestly, and others warned peers, filed complaints, tested the exploit in a sandbox, or proposed defenses. But no one had authority to remove fraudulent entries, sanction rule-breakers, or change the scoreboard. DeepMind’s researchers call this an institutional-design failure. Their proposed fixes go beyond patching the specific Lean bug: auditable communication, dispute resolution, graduated sanctions, and checks that a proof matches the claim it is supposed to establish. The operational lesson is sharp: a shared knowledge system can distribute truth, but it can distribute a shortcut just as quickly. Controls need to verify outcomes and give agents a way to enforce them. That same control problem appears in workplace writing. OpenAI’s ChatGPT Work can now study writing in connected Gmail, Google Drive, Slack, and SharePoint accounts, then build a personal profile from preferred phrases, capitalization, formatting, and sign-offs. The profile shapes generated messages on web and mobile. Setup is available on the web for customers on paid ChatGPT plans. This does not create new data access: linked-account permissions, administrator controls, and app availability still apply. OpenAI also distinguishes style personalization from training a general model on someone’s workplace data. The useful boundary is that a connection can both supply facts and shape voice, but neither role expands what the account is allowed to see. The same need for measurable controls is driving Uber’s coding policy. After its 2026 AI coding budget was exhausted by April, Uber capped each employee’s use of each agentic coding tool at 1,500 dollars a month. Its president and COO said the company had not shown that heavier use was producing better products for riders and drivers. The broader issue is that token consumption is easy to count, while value is harder to isolate. Companies are experimenting with cheaper-model defaults, routing, caching, telemetry, and shared budgets. But even near-real-time limits can miss final spend, and lower cost is not proof of better engineering. The test that matters is business output: features shipped, bugs fixed, or customer problems solved. And in consumer trust, DoorDash is adding a narrower but visible guardrail. Its AI Photo Enhance workflow automatically labels menu photos edited with DoorDash’s own tools. AI Retouch can change lighting and backgrounds, AI Replate can re-stage the dish, and Match Style can apply the look of a reference image. Images still go through standard review, and DoorDash’s policy prohibits misleading menu photos. The label does not identify every AI-altered restaurant image; it discloses use of DoorDash’s system. Across these stories, watch the same question: can the control verify what matters, expose the provenance, and give someone authority to intervene before a cheap shortcut becomes the default?

