Anthropic opens safety reviews to outside evaluators
Dario Amodei also wants slower capability gains, but shared standards still require other labs and governments to opt in.
By Saeed Ezzati8 min read

The audio edition
Listen to this newsletter
0:003:19
Read transcript
The clearest operating change this week comes from Anthropic. CEO Dario Amodei says independent safety evaluators will now get employee-level access to the company’s safety work—the same access available to internal risk-assessment teams. They can examine practices, report incidents, and publish findings without Anthropic’s editorial control. Amodei says the arrangement takes effect immediately, turning outside review into a standing mechanism inside one frontier AI company. His larger proposal is much harder to implement. Amodei is calling for developers to slow capability gains, use the extra time to improve safety and judgment, establish permanent independent evaluation across frontier companies, and agree on common standards among democratic countries. He also wants international coordination on shared risks, including a ban on using AI to develop biological weapons. Those broader steps remain voluntary proposals, not industry rules. The argument for slowing down is that increasingly capable systems might help create their successors, producing recursive self-improvement that moves faster than people can understand or control. Amodei is not calling for progress to stop; he is asking companies to pace the frontier carefully. The important distinction is between what Anthropic controls directly—access and publication rights for evaluators—and what depends on competitors and governments choosing the same restraint. The test now is whether independent review produces findings that materially change decisions, and whether anyone else adopts the model. That split between voluntary restraint and measurable rules also appears in OpenAI’s financing decision. Sam Altman says OpenAI will not pursue an IPO in 2026, calling the timing ill-advised while safety, alignment, business readiness, and society’s response to more capable AI remain unresolved. He left later timing open. OpenAI has discussed pauses at new capability levels and coordination with other companies and governments, but gave no specific readiness test, safeguards, or development freeze. The useful question is what evidence eventually turns delay into a decision. The same gap between long-range promise and current evidence shapes Arm CEO Rene Haas’s outlook. He says AI could eventually help cure cancer, while acknowledging that today’s systems cannot model how cancer affects a DNA marker. The nearer-term result is narrower: the NHS says AI-powered X-ray tools helped more than four million patients receive faster lung diagnoses earlier this year. Haas also expects humanoid robots to spread within five years, but says chip shortages are holding deployment back. Watch demonstrated gains in difficult biology, not just forecasts. And at the builder level, OpenAI’s Codex guidance makes the same point operational. For GPT-6 Astra, teams should trim broad skill descriptions, turn AGENTS.md files into conditional maps, and define exactly what “done” means. Safe, reversible local work can proceed, while production access, destructive migrations, external effects, and credentials retain approval gates. Astra may stop early when a workflow lacks a clear end state. Across these stories, capability is moving faster than shared operating rules. The practical thing to watch is whether companies replace broad assurances with independent findings, explicit thresholds, and controls tied to real-world risk.

