Anthropic opens safety reviews to outside evaluators

Dario Amodei also wants slower capability gains, but shared standards still require other labs and governments to opt in.

By 8 min read
Anthropic opens safety reviews to outside evaluators
Anthropic opens safety reviews to outside evaluators

The audio edition

Listen to this newsletter

About 3:19
0:003:19
Read transcript
The clearest operating change this week comes from Anthropic. CEO Dario Amodei says independent safety evaluators will now get employee-level access to the company’s safety work—the same access available to internal risk-assessment teams. They can examine practices, report incidents, and publish findings without Anthropic’s editorial control. Amodei says the arrangement takes effect immediately, turning outside review into a standing mechanism inside one frontier AI company. His larger proposal is much harder to implement. Amodei is calling for developers to slow capability gains, use the extra time to improve safety and judgment, establish permanent independent evaluation across frontier companies, and agree on common standards among democratic countries. He also wants international coordination on shared risks, including a ban on using AI to develop biological weapons. Those broader steps remain voluntary proposals, not industry rules. The argument for slowing down is that increasingly capable systems might help create their successors, producing recursive self-improvement that moves faster than people can understand or control. Amodei is not calling for progress to stop; he is asking companies to pace the frontier carefully. The important distinction is between what Anthropic controls directly—access and publication rights for evaluators—and what depends on competitors and governments choosing the same restraint. The test now is whether independent review produces findings that materially change decisions, and whether anyone else adopts the model. That split between voluntary restraint and measurable rules also appears in OpenAI’s financing decision. Sam Altman says OpenAI will not pursue an IPO in 2026, calling the timing ill-advised while safety, alignment, business readiness, and society’s response to more capable AI remain unresolved. He left later timing open. OpenAI has discussed pauses at new capability levels and coordination with other companies and governments, but gave no specific readiness test, safeguards, or development freeze. The useful question is what evidence eventually turns delay into a decision. The same gap between long-range promise and current evidence shapes Arm CEO Rene Haas’s outlook. He says AI could eventually help cure cancer, while acknowledging that today’s systems cannot model how cancer affects a DNA marker. The nearer-term result is narrower: the NHS says AI-powered X-ray tools helped more than four million patients receive faster lung diagnoses earlier this year. Haas also expects humanoid robots to spread within five years, but says chip shortages are holding deployment back. Watch demonstrated gains in difficult biology, not just forecasts. And at the builder level, OpenAI’s Codex guidance makes the same point operational. For GPT-6 Astra, teams should trim broad skill descriptions, turn AGENTS.md files into conditional maps, and define exactly what “done” means. Safe, reversible local work can proceed, while production access, destructive migrations, external effects, and credentials retain approval gates. Astra may stop early when a workflow lacks a clear end state. Across these stories, capability is moving faster than shared operating rules. The practical thing to watch is whether companies replace broad assurances with independent findings, explicit thresholds, and controls tied to real-world risk.
Dario Amodei also wants slower capability gains, but shared standards still require other labs and governments to opt in.
Weekly digest / The Weekly Digest Sunday, September 13, 2026
Our tools Superpower ChatGPT/WFH.team/Snipman

This week's briefing

What happened this week

Inside this week's digest
01
02
03
04
Anthropic’s Dario Amodei Calls for Slower AI Progress and Outside Safety Reviews

Lead story / security risk

Anthropic gives outside reviewers access to its safety work

Read full story  ↗
A tool for your workflow WFH.team A focused feed of carefully selected remote roles and practical work-from-home resources. Remote work, without the noisy job-board scroll
 
Browse remote roles  ↗
OpenAI Says It Will Skip a 2026 IPO Over AI Safety Concerns

business

OpenAI rules out a 2026 IPO over safety concerns

Continue reading  ↗
Arm CEO Says AI Could Help Cure Cancer Within His Lifetime

business

Arm CEO says AI could help cure cancer, but current systems fall short

Continue reading  ↗
OpenAI Publishes a Leaner Prompting Playbook for Codex Agents

tools

OpenAI tells developers to use leaner Codex prompts

Continue reading  ↗
Pentagon Says It Has Moved 90% of Classified AI Work Off Anthropic

government action

Pentagon moves 90% of classified AI work off Anthropic

Continue reading  ↗
The Weekly Digest themed section header

Who tests powerful models, whether competitors can coordinate on safety, and what a voluntary slowdown would require.

Senate AI Bill Hits Dispute Over Who Tests Powerful ModelsRead story ↗
OpenAI Seeks Congress’s View on a Coordinated AI SlowdownRead story ↗
Sam Altman Reportedly Told Staff OpenAI Could Slow AI-Agent Development With Other LabsRead story ↗
 

Weekend tool drop

5 AI tools worth knowing this weekend

Selected for fit, not rank
QApilot MCP for Android Runs plain-English Android tests and saves passes as replayable Gherkin cases. Best for / Android teams using coding agents Open ↗
Youkti Tracks accounts and deals to recommend the next sales action from signals and history. Best for / AEs, RevOps, and outbound teams Open ↗
TIM PG Masks clipboard and document data locally before sending text to LLMs. Best for / Privacy-conscious Windows AI users Open ↗
Suno v6 Creates and edits music from text, audio, images, or video inputs. Best for / Creators iterating on AI music Open ↗
Live Captions by Subanana Delivers live translated captions or audio to audience phones, screens, and broadcast feeds. Best for / Multilingual events and presentations Open ↗
 
Meta’s Muse Reaches No. 2 on U.S. App Store, Sensor Tower Estimates Meta’s Muse reaches No. 2 in the U.S. App Store ↗launch
Roblox Adds AI Game Tools and Plans Browser and Standalone Game Access Roblox expands AI game tools and plans wider distribution ↗launch
OpenAI Halts New $200 Pro Sign-Ups as Astra Demand Strains Capacity OpenAI pauses new $200 Pro sign-ups over Astra demand ↗platform shift
Apple Introduces iPhone Duo With an AI-Assisted Hinge Apple uses AI in iPhone Duo hinge production ↗launch
 

The Internet Had a Point

 
From our network. Tools built for the way you work. Useful products from the team behind Superpower Daily.

Reader check-in

Help shape tomorrow's briefing

One click tells us what to keep, improve, or tighten.

Superpower Daily tracks the companies, models, products, tools, policy decisions, and cultural shifts moving AI. Follow Superpower Daily Email preferences